Skip to content

Storage recovery

Durable recovery patterns for replicated block storage. The exact commands live in your platform repo; this page is the judgement — especially the two ways a recovery silently loses data.

Stale VolumeAttachment after a node reboot

After a node reboot, Talos upgrade, or instance-manager recovery, single-replica volumes commonly wedge: pods stick in Init/ContainerCreating with "volume hasn't been attached yet," while the VolumeAttachment says attached=true and the storage layer says the volume is detached. The CSI attacher sees attached=true and never re-triggers, but the engine actually stopped at reboot. Delete the stale VolumeAttachment — the attacher recreates it and the volume re-attaches within seconds. A related flavour ("staging target path no longer valid," an unresolved upstream kubelet issue) has the same fix — deleting the stale VA forces a clean detach/re-attach; restarting the CSI plugin pod does not, because the stale state is on the kubelet's disk, not in CSI memory.

Before salvaging from a surviving replica, check its last-healthy timestamp

A salvage from a stale replica loses data silently

Recovering a volume from its one surviving replica (a UI "Salvage" or activating a replica) brings it back attached and I/O-clean even if that replica is stale — everything written after the replica's last-healthy time is silently gone. Capture the replicas' last-healthy timestamps first, before activating anything: once a replica re-attaches, that field updates to now and the staleness is no longer visible. Pick the freshest replica; if even the freshest is meaningfully stale, prefer restoring from backup — a day-old backup can be fresher than a stale replica — and stop to consider whether the application can tolerate the loss before attaching.

When a salvage stalls, stop patching

Over-intervention corrupts the volume permanently

A single-replica salvage that isn't progressing is not fixed by piling on manual patches (clearing failure timestamps, deleting engines repeatedly, patching status, restarting managers) — especially while the environment is still unhealthy. That drives the volume's custom resource into a terminal state from which nothing recovers. When a salvage stalls: verify environment health first (the target engine image fully deployed on every node, nodes stable, no manager churn), then retry only the minimal clean recipe. Manager restarts are not a free diagnostic — they flap the control plane.

After a forced re-attach, expect embedded caches to need rebuilding

A volume that was force-reattached can leave an application holding an embedded database (LMDB/SQLite and the like) on it in a torn state. Expect it, and re-derive that cache from its authoritative source rather than trying to repair it in place.

See also: Storage · Node upgrade & maintenance · Database operations & recovery.