> ## Documentation Index
> Fetch the complete documentation index at: https://private-7c7dfe99-parallel-read-in-order-multi-part.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# How to recover from a corrupt Keeper snapshot

> Article describing how to recover from a corrupt Keeper snapshot: how the problem manifests, what a snapshot is and where to find it and possible recovery strategies.

Corrupt or bad ClickHouse Keeper snapshots can cause significant system instability, such as metadata inconsistencies, read-only states for tables, resource exhaustion, or failed backups. This article covers:

* [What snapshots are and where to find them](#overview)
* [How the problem manifests](#symptoms)
* [Possible strategies for recovery](#recovery-strategies) and what each of them means

<h2 id="overview">
  Overview of Keeper snapshots
</h2>

<h3 id="what-is-snapshot">
  What is a snapshot?
</h3>

A snapshot is a serialized state of Keeper's internal data (such as metadata about clusters, table coordination paths, and configurations) at a specific point in time. Snapshots are vital for resynchronizing Keeper nodes within a cluster, recovering metadata during failures, and start-up or restart processes that rely on a known-good Keeper state.

<h3 id="where-to-find-snapshots">
  Where can I find snapshots?
</h3>

Snapshots are stored as files on the local filesystem of Keeper nodes. By default, they are stored at `/var/lib/clickhouse/coordination/snapshots/` or by the custom path specified by `snapshot_storage_path` in your `keeper_server.xml` file. Snapshots are named incrementally (e.g., snapshot.23), with newer ones having higher numbers.

For multi-node clusters, each Keeper node has its own snapshot directory.

<Note>
  Consistency within snapshots across nodes is critical for recovery.
</Note>

<h2 id="symptoms">
  Key symptoms and manifestations of corrupt Keeper snapshots
</h2>

The table below details some common symptoms and manifestations of corrupt Keeper snapshots:

| **Category** | **Issue Type** | **What to look for** |
| - | - | - |
| **Operational Issues** | Read-Only Mode | Tables unexpectedly switch to read-only mode |
| | Query Failures | Persistent query failures with `Coordination::Exception` errors |
| **Metadata Corruption** | Outdated Metadata | Dropped tables not reflected; operation failures due to stale metadata |
| **Resource Overload** | System Resource Exhaustion | Keeper nodes consume excessive CPU, memory, or disk space; potential downtime |
| | Disk Full | Disk full during snapshot creation |
| **Backup & Restore** | Backup Failures | Backups fail due to missing or inconsistent Keeper metadata |
| **Snapshot Creation/Transfer** | Keeper Crash | Keeper crash mid-snapshot (look for "SEGFAULT" errors) |
| | Snapshot Transfer Corruption | Corruption during snapshot transfer between replicas |
| | Race Condition | Race condition during log compaction - background commit thread accessing deleted logs |
| | Network Synchronization | Network issues preventing snapshot sync from leader to followers |

**Log Indicators:**

Before diagnosing snapshot corruption, check **Keeper logs** for specific error patterns:

| **Log Type** | **What to Look For** |
| - | - |
| **Snapshot corruption errors** | • `Aborting because of failure to load from latest snapshot with index`<br />• `Failure to load from latest snapshot with index {}: {}. Manual intervention is necessary for recovery`<br />• `Failed to preprocess stored log at index {}, aborting to avoid inconsistent state`<br />• Snapshot serialization/loading failures during startup |
| **Other Keeper issues** | • `Coordination::Exception`<br />• `Zookeeper::Session Timeout`<br />• Synchronization or election issues<br />• Log compaction race conditions |

<h2 id="recovery-strategies">
  Recovering from corrupt Keeper snapshots
</h2>

Before touching any files, always:

1. Stop all Keeper nodes to prevent further corruption
2. Backup everything by copying the entire coordination directory to a safe location
3. Verify cluster quorum to ensure at least one node has good data

***

<h3 id="restore-from-existing-backup">
  1. Restore from an existing backup
</h3>

You should follow this process if:

* The Keeper metadata or snapshot corruption makes current data unsalvageable.
* A backup exists with a known-good Keeper state.

Follow the steps below to restore an existing backup:

1. Locate and validate the newest backup for metadata consistency.
2. Shut down the ClickHouse and Keeper services.
3. Replace the faulty snapshots and logs with those from the backup directory.
4. Restart the Keeper cluster and validate metadata synchronization.

<Tip>
  **Backup regularly**

  If backups are outdated, you may incur a loss of recent metadata changes. For this reason, we recommend backing up regularly.
</Tip>

***

<h3 id="rollback-to-older-snapshot">
  2. Rollback to an older snapshot
</h3>

You should follow this process when:

* Recent snapshots are corrupt, but older ones remain usable.
* Incremental logs are intact for consistent recovery.

Follow the steps below to roll back to an older snapshot:

1. Identify and select a valid older snapshot (e.g., snapshot.19) from the Keeper directory.
2. Remove newer snapshots and logs.
3. Restart Keeper so it replays logs to rebuild the metadata state.

<Warning>
  **Metadata desynchronization risk**

  There is a risk of metadata desynchronization if snapshots and logs are missing or incomplete.
</Warning>

***

<h3 id="restore-metadata-with-system-restore-replica">
  3. Restore metadata using `SYSTEM RESTORE REPLICA`
</h3>

You should follow this process when:

* Keeper metadata is lost or corrupted but table data still exists on disk
* Tables have switched to read-only mode due to missing ZooKeeper/Keeper metadata
* You need to recreate metadata in Keeper based on locally available data parts

Follow the steps below to restore metadata:

1. Verify that table data exists locally in your clickhouse-server data path, set by `<path>` in your config. (`/var/lib/clickhouse/data/` by default)

2. For each affected table, execute:

```sql theme={null}
SYSTEM RESTART REPLICA [db.]table_name;
SYSTEM RESTORE REPLICA [db.]table_name;
```

3. For database-level recovery (if using Replicated database engine):

```sql theme={null}
SYSTEM RESTORE DATABASE REPLICA db_name;
```

4. Wait for synchronization to complete:

```sql theme={null}
SYSTEM SYNC REPLICA [db.]table_name;
```

5. Verify recovery by checking `system.replicas` for `is_readonly = 0` and monitoring `system.detached_parts`

<Info>
  **How it works**

  `SYSTEM RESTORE REPLICA` detaches all existing parts, recreates metadata in Keeper (as if it's a new empty table), then reattaches all parts. This avoids re-downloading data over the network.
</Info>

<Warning>
  **Prerequisites**

  This only works if local data parts are intact. If data is also corrupted, use strategy #5 (rebuild cluster) instead.
</Warning>

***

<h3 id="drop-and-recreate-replica-metadata">
  4. Drop and recreate replica metadata in Keeper
</h3>

You should follow this process when:

* The error occurs on a single replica of the cluster and has corrupt or inconsistent metadata in Keeper
* You encounter errors like "Part XXXXX intersects previous part YYYYY"
* You need to completely reset a replica's Keeper metadata while preserving local data

Follow the steps below to drop and recreate metadata:

1. On the affected replica, detach the table:

```sql theme={null}
DETACH TABLE [db.]table_name;
```

2. Remove the replica's metadata from Keeper (execute on any replica):

```sql theme={null}
SYSTEM DROP REPLICA 'replica_name' FROM ZKPATH '/clickhouse/tables/{shard}/table_name';
```

To find the correct ZooKeeper path:

```sql theme={null}
SELECT zookeeper_path, replica_name FROM system.replicas WHERE table = 'table_name';
```

3. Reattach the table (it will be in read-only mode):

```sql theme={null}
ATTACH TABLE [db.]table_name;
```

4. Restore the replica metadata:

```sql theme={null}
SYSTEM RESTORE REPLICA [db.]table_name;
```

5. Synchronize with other replicas:

```sql theme={null}
SYSTEM SYNC REPLICA [db.]table_name;
```

6. Check `system.detached_parts` on all replicas after recovery

<Warning>
  **Execute on all affected replicas**

  If the corruption affects multiple replicas, repeat these steps on each one sequentially.
</Warning>

<Tip>
  **For entire database**

  If using a Replicated database, you can use `SYSTEM DROP REPLICA ... FROM DATABASE db_name` instead.
</Tip>

**Alternative: Using force\_restore\_data flag**

For automatic recovery of all replicated tables at server startup:

1. Stop ClickHouse server
2. Create the recovery flag:

```bash theme={null}
sudo -u clickhouse touch /var/lib/clickhouse/flags/force_restore_data
```

3. Start ClickHouse server
4. The server will automatically delete the flag and restore all replicated tables
5. Monitor logs for recovery progress

This approach is useful when multiple tables need recovery simultaneously.

***

<h3 id="rebuild-keeper-cluster">
  5. Rebuild Keeper cluster
</h3>

You should follow this process when:

* No valid snapshots, logs, or backups are available for recovery.
* You need to recreate the entire Keeper cluster and its metadata.

Follow the steps below to rebuild the Keeper cluster:

1. Fully stop the ClickHouse and Keeper clusters.
2. Reset each Keeper node by cleaning the snapshot and log directories.
3. Initialize one Keeper node as the leader and add other nodes incrementally.
4. Re-import metadata if available from external records.

<Warning>
  **Time-intensive process**

  This process is time-intensive and carries a risk of prolonged outage. Total data reconstruction is required.
</Warning>

***

<h3 id="remove-orphaned-nodes-on-startup">
  6. Remove orphaned nodes from the snapshot on startup
</h3>

An *orphaned* node is a node whose parent is missing from the snapshot. Keeper refuses to load such a
snapshot, logging `Failure to load from latest snapshot with index`.

You should follow this process when:

* The snapshot is structurally damaged in exactly this way, and
* No valid backup or older snapshot is available (strategies 1 and 2 are preferable whenever they apply).

Set both of the following under `<keeper_server>` and restart the node:

```xml theme={null}
<keeper_server>
    <remove_orphaned_nodes_on_startup>true</remove_orphaned_nodes_on_startup>
    <digest_enabled>false</digest_enabled>
</keeper_server>
```

<Note>
  **Requires the in-memory nodes storage**

  The cleanup is implemented only by the in-memory nodes storage. If `use_lsmt_storage` is enabled under
  `<coordination_settings>`, Keeper refuses to start with `remove_orphaned_nodes_on_startup` set: disable
  `use_lsmt_storage` for the recovery restart, then re-enable it after a successful startup.
</Note>

Keeper then drops the orphaned subtrees and starts. The removal is reported in the log at `WARNING`
level as two lines: one with the total number of removed nodes and the roots of the removed orphaned
subtrees, and one with the paths missing from the snapshot that bound the damaged region. Descendants
below those roots are removed as well but are not listed individually.

After checking the local log tail, Keeper writes the repaired tree to a new snapshot file at the same
log index before starting Raft. Completed repairs take precedence over older copies at that index,
so an interrupted recovery can resume without selecting the damaged snapshot. New followers receive
the repaired snapshot, and subsequent restarts
no longer need orphan removal enabled. The original snapshot and any same-index copies are preserved
with an `orphaned_` filename prefix; these files are excluded from snapshot loading and automatic
retention. Keep them for diagnosis, then remove them manually when they are no longer needed. The
recovery restart needs additional disk space for the repaired snapshot and the backup of the original.
If writing the repaired snapshot, synchronizing local directory changes, or preserving the originals
fails, Keeper refuses to start. A subsequent startup finishes any interrupted archival.

<Warning>
  **Repairs local state only**

  This removes nodes from *this* node's copy of the data. Peers that still hold those nodes will diverge
  from it, and with `digest_enabled` set to `false` nothing detects that divergence afterwards. Treat this
  as single-node recovery, or apply it to every node in the cluster. Set both settings back to their
  defaults once the node is up.
</Warning>

Keeper deliberately refuses to start if a local log entry above the snapshot references the damaged
region: a removed node, anything below it, or the direct parent of a removed subtree. The direct parent
counts as a conflict because its set of children no longer matches the tree the log was written against,
so even an entry that only touches it — including a create or remove of a *sibling* under it, which
updates its repaired child counters — could replay differently. Entries that look only at the parent's
existence, version or modification ids, such as a `Check`, a `CheckNotExists`, a `CheckStat` that compares neither `numChildren` nor `cversion`,
a `Set` or `SetACL` that rewrites only the parent's own data or ACL, or
the data, child and exist watches a client re-registers after a reconnect, are unaffected and do not
block recovery; persistent watches re-registered by a client never consult the tree and never block it.
Nodes further up the chain are not
affected: they keep exactly the data, version and children they have on a healthy peer, so an entry that
only touches such a node replays identically and does not block recovery. The exception is an entry that
walks the whole subtree below its path, such as a recursive removal or a recursive listing: those do
observe the removed nodes and conflict at any height, unless Keeper rejects the request before looking
at the tree (for example, a recursive removal of `/`). The same applies to a session close
(`Close`) in the log tail when that session owned a removed ephemeral node, or owns a surviving ephemeral
node directly under such a parent: closing the session removes its ephemeral nodes and updates their
parents' counters. Replaying any such entry against the
repaired tree would silently produce a different state. If that happens, recover the node from a healthy
peer instead: stop it, clear its coordination directory, and let it re-sync from the leader, or use
strategy 1 or 2.
