Storage recovery¶
Durable recovery patterns for replicated block storage. The exact commands live in your platform repo; this page is the judgement — especially the two ways a recovery silently loses data.
Stale VolumeAttachment after a node reboot¶
After a node reboot, Talos upgrade, or instance-manager recovery, single-replica volumes
commonly wedge: pods stick in Init/ContainerCreating with "volume hasn't been attached
yet," while the VolumeAttachment says attached=true and the storage layer says the volume is
detached. The CSI attacher sees attached=true and never re-triggers, but the engine
actually stopped at reboot. Delete the stale VolumeAttachment — the attacher recreates it
and the volume re-attaches within seconds. A related flavour ("staging target path no longer
valid," an unresolved upstream kubelet issue) has the same fix — deleting the stale VA
forces a clean detach/re-attach; restarting the CSI plugin pod does not, because the stale
state is on the kubelet's disk, not in CSI memory.
Before salvaging from a surviving replica, check its last-healthy timestamp¶
A salvage from a stale replica loses data silently
Recovering a volume from its one surviving replica (a UI "Salvage" or activating a replica) brings it back attached and I/O-clean even if that replica is stale — everything written after the replica's last-healthy time is silently gone. Capture the replicas' last-healthy timestamps first, before activating anything: once a replica re-attaches, that field updates to now and the staleness is no longer visible. Pick the freshest replica; if even the freshest is meaningfully stale, prefer restoring from backup — a day-old backup can be fresher than a stale replica — and stop to consider whether the application can tolerate the loss before attaching.
When a salvage stalls, stop patching¶
Over-intervention corrupts the volume permanently
A single-replica salvage that isn't progressing is not fixed by piling on manual patches (clearing failure timestamps, deleting engines repeatedly, patching status, restarting managers) — especially while the environment is still unhealthy. That drives the volume's custom resource into a terminal state from which nothing recovers. When a salvage stalls: verify environment health first (the target engine image fully deployed on every node, nodes stable, no manager churn), then retry only the minimal clean recipe. Manager restarts are not a free diagnostic — they flap the control plane.
After a forced re-attach, expect embedded caches to need rebuilding¶
A volume that was force-reattached can leave an application holding an embedded database (LMDB/SQLite and the like) on it in a torn state. Expect it, and re-derive that cache from its authoritative source rather than trying to repair it in place.
See also: Storage · Node upgrade & maintenance · Database operations & recovery.