Node upgrade & maintenance¶
Talos is API-managed and immutable — an upgrade replaces the OS image rather than patching in place. This page is the operational discipline; the exact commands and installer-image schematics live in the talos-cluster README (sections Rolling upgrades, Applying patch changes, Per-node etcd metrics patches).
Golden rules¶
- One node at a time. Never script or loop a destructive rollout across nodes — a bad image or a wedged reboot must stop at one node, not cascade. Drive each node by hand and verify before moving on.
- Talos and Kubernetes minors move together. A Talos upgrade carries a new kubelet, so treat it as a Kubernetes minor too: hold at the current Talos minor until every workload supports the next kubelet version. Clear the compatibility gate first, then upgrade nodes.
- Privileged Talos operations (
talosctl upgrade/apply-config/patch) reboot or rewrite a node — they are run deliberately by an operator, never automated.
Before touching any node¶
- Storage is healthy — all Longhorn volumes attached and healthy, no replica rebuild in flight. Rebooting a node mid-rebuild risks the volume (see Storage).
- Cluster is green — nodes
Ready, no pending critical alerts.
Workers¶
Use the repo's safe-reboot.sh <node> <worker-installer-image> — it drains, evicts, runs
talosctl upgrade --preserve, waits for the node to return, and uncordons.
Control-plane nodes¶
Never run safe-reboot.sh on a control-plane node (it is worker-only by design). Instead,
one CP at a time: talosctl upgrade with the control-plane installer image, reboot, then
verify etcd quorum + API health and wait ~60 s to settle before the next node. Note two
CP-specific traps documented in the README:
- Some settings (e.g.
cluster.etcd.*) only take effect on reboot. - A full
apply-configof a regeneratedcontrolplane.yamlwipes per-node patches (e.g. the etcd-metrics patches) — re-apply them with the same reboot, or use the README's pre-merge variant.
Draining a node that hosts single-replica volumes¶
Multi-replica volumes migrate off a draining node cleanly. Single-replica strict-local volumes cannot migrate — their one copy is pinned to the node — so a naïve drain hangs, and a forceful drain corrupts.
- Never drain with eviction disabled. A raw-delete drain (
--disable-eviction) bypasses every PodDisruptionBudget, including the storage manager's, and can kill the node's instance-manager while volumes are still attached — every engine on the node dies at once, producing a fault storm. Use the eviction API so pods leave in order: workloads evict → their volumes detach → replicas stop → the manager's PDB releases → the manager evicts last with nothing attached. - Pair it with a replica-stopped drain policy. Set the storage layer's node-drain policy to "allow if replica is stopped" so it enforces the ordering through the manager PDB. The policy and the eviction-API drain are two halves of one mechanism — either alone is inert.
- Strict-local pods still wedge the drain, because their StatefulSet pod re-creates
Pending(pinned to the cordoned node) and keeps the PVC claimed, so the volume never detaches. Once the ordinary workloads have evicted, delete the node's remaining Longhorn VolumeAttachments so the idle volumes detach and the manager can finish leaving.
Reboot fast enough to earn a delta sync, not a full rebuild
Longhorn gives a replica that returns within its replenishment-wait interval a cheap delta resync; a replica that returns after it gets a full rebuild. The down-clock starts when the volume detaches (just before reboot), so a normal reboot easily fits the window — but a node left down for a long maintenance action will pay the full-rebuild cost on every replica it hosted. (A few volume types disable the revision counter for write throughput and always full-rebuild regardless — harmless, just slower.)
Extra pre-reboot gates when the node hosts a database¶
Beyond "storage healthy, cluster green," clear these before rebooting a node that hosts a clustered database:
- Galera (MariaDB): confirm the cluster is fully synced (all members
Synced, full membership) and that a state transfer could actually run, before the reboot — and wait for it to re-sync fully after, before moving to the next node. A Galera StatefulSet showingreplicas: 0is usually the operator mid-recovery (it scales down to pick a bootstrap node by sequence number, then scales back), not a failure — leave it alone. - Primary/replica databases (e.g. Postgres): reboot a node holding only a replica freely; reboot the node holding the primary last, so automatic failover relocates it once, cleanly.
After the upgrade¶
- Bump the local tooling to match the cluster:
talosctl, andlonghornctlafter a Longhorn bump. - Update the version references in the repo.
See also: Storage · Operations · Architecture · Storage recovery · Database operations & recovery.