Storage¶
Persistent storage is provided by Longhorn (CSI provisioner
driver.longhorn.io) — replicated block storage backed by the nodes' local disks.
StorageClasses¶
| Class | Encrypted | Access | Binding |
|---|---|---|---|
longhorn (default) |
no | RWO | Immediate |
longhorn-static |
no | RWO | Immediate |
longhorn-local |
no | RWO | WaitForFirstConsumer (node-local) |
longhorn-rwx |
no | RWX | Immediate |
longhorn-crypto |
yes | RWO | Immediate |
longhorn-crypto-local |
yes | RWO | WaitForFirstConsumer |
longhorn-crypto-rwx |
yes | RWX | Immediate |
-cryptoclasses encrypt the volume at rest (per-volume key).-rwxclasses are ReadWriteMany (shared), via Longhorn's share-manager.-local/WaitForFirstConsumerbinds the volume on the node where the first consumer pod schedules.- All classes allow volume expansion — grow the PVC, then the filesystem resizes (a pod restart picks up the new size).
Replica posture¶
A StorageClass's replica count decides whether a node or disk loss self-heals or requires a restore:
- Multi-replica classes (the default) keep copies on several nodes — a node or disk loss is repaired by an automatic rebuild, no data lost.
- Single-replica
strict-localclasses keep exactly one copy, pinned to the node the volume was created on. They cannot migrate and cannot rebuild — a lost node or disk means restore from backup, full stop. Their durability rests entirely on the backup tracks.
Databases and other latency-sensitive state are often placed on single-replica strict-local volumes deliberately (local disk, no replication overhead). That is a sound choice only if those volumes are backed up and you accept restore-not-rebuild for them. Know which of your volumes are single-replica before you reboot a node — see Node upgrade & maintenance.
Backups¶
Volume and database backups run on a schedule to external, S3-compatible object storage. The backup layout, restore procedures, and disaster-recovery runbooks are maintained in the private operations documentation.
Longhorn's recurring-job label lives on the Volume, not the PVC
To opt a volume into a Longhorn recurring backup job, the
recurring-job.longhorn.io/<job>: enabled label must be on the volume.longhorn.io
custom resource, not on the PVC. Labelling the PVC is a silent no-op — the volume
is simply never backed up, with no error anywhere. Verify coverage by checking the label
on the Volume CR and confirming backups actually exist for it, not by assuming the PVC
label took effect.
Longhorn upgrades¶
Treat every Longhorn upgrade as a maintenance window — never auto-merged
Any Longhorn version change (patch releases included) rolls the storage-manager DaemonSet
and triggers a single-replica engine-image migration. The migration's reconcile spikes
manager memory regardless of how small the version jump is, and single-replica attached
volumes are skipped by the automatic engine upgrade — they need a manual spec.image
patch per volume afterwards. Land Longhorn bumps by hand, in a window, and watch the roll.
Do not let dependency automation merge them (see Renovate).
Storage on small nodes¶
On memory-constrained worker nodes, a few Longhorn settings are protective, not performance knobs — do not raise them to "go faster":
- Concurrent-rebuild limits (per-node and per-volume) cap how many replica rebuilds run at once. Parallel rebuild storms can OOM-kill the storage manager. Expect rebuilds after a reboot to take time; that is the cost of safety. The real remedy for slowness is more RAM, not more concurrency.
- Give the storage manager a generous memory limit. A container memory limit is not a reservation — only requests count against a node's allocatable memory; the limit is just the burst ceiling. A manager that idles at a couple hundred MiB can carry a multi-GiB limit at zero steady-state cost, and that headroom is exactly what keeps the upgrade cache-sync spike from being cgroup-OOM-killed. Keep the request modest, the limit roomy, and give the manager a high-priority PriorityClass so it is evicted last under node pressure.
See also: Operations · Architecture · Node upgrade & maintenance · Capacity & resource limits.