Skip to content

Node upgrade & maintenance

Talos is API-managed and immutable — an upgrade replaces the OS image rather than patching in place. This page is the operational discipline; the exact commands and installer-image schematics live in the talos-cluster README (sections Rolling upgrades, Applying patch changes, Per-node etcd metrics patches).

Golden rules

  • One node at a time. Never script or loop a destructive rollout across nodes — a bad image or a wedged reboot must stop at one node, not cascade. Drive each node by hand and verify before moving on.
  • Talos and Kubernetes minors move together. A Talos upgrade carries a new kubelet, so treat it as a Kubernetes minor too: hold at the current Talos minor until every workload supports the next kubelet version. Clear the compatibility gate first, then upgrade nodes.
  • Privileged Talos operations (talosctl upgrade / apply-config / patch) reboot or rewrite a node — they are run deliberately by an operator, never automated.

Before touching any node

  • Storage is healthy — all Longhorn volumes attached and healthy, no replica rebuild in flight. Rebooting a node mid-rebuild risks the volume (see Storage).
  • Cluster is green — nodes Ready, no pending critical alerts.

Workers

Use the repo's safe-reboot.sh <node> <worker-installer-image> — it drains, evicts, runs talosctl upgrade --preserve, waits for the node to return, and uncordons.

Control-plane nodes

Never run safe-reboot.sh on a control-plane node (it is worker-only by design). Instead, one CP at a time: talosctl upgrade with the control-plane installer image, reboot, then verify etcd quorum + API health and wait ~60 s to settle before the next node. Note two CP-specific traps documented in the README:

  • Some settings (e.g. cluster.etcd.*) only take effect on reboot.
  • A full apply-config of a regenerated controlplane.yaml wipes per-node patches (e.g. the etcd-metrics patches) — re-apply them with the same reboot, or use the README's pre-merge variant.

Draining a node that hosts single-replica volumes

Multi-replica volumes migrate off a draining node cleanly. Single-replica strict-local volumes cannot migrate — their one copy is pinned to the node — so a naïve drain hangs, and a forceful drain corrupts.

  • Never drain with eviction disabled. A raw-delete drain (--disable-eviction) bypasses every PodDisruptionBudget, including the storage manager's, and can kill the node's instance-manager while volumes are still attached — every engine on the node dies at once, producing a fault storm. Use the eviction API so pods leave in order: workloads evict → their volumes detach → replicas stop → the manager's PDB releases → the manager evicts last with nothing attached.
  • Pair it with a replica-stopped drain policy. Set the storage layer's node-drain policy to "allow if replica is stopped" so it enforces the ordering through the manager PDB. The policy and the eviction-API drain are two halves of one mechanism — either alone is inert.
  • Strict-local pods still wedge the drain, because their StatefulSet pod re-creates Pending (pinned to the cordoned node) and keeps the PVC claimed, so the volume never detaches. Once the ordinary workloads have evicted, delete the node's remaining Longhorn VolumeAttachments so the idle volumes detach and the manager can finish leaving.

Reboot fast enough to earn a delta sync, not a full rebuild

Longhorn gives a replica that returns within its replenishment-wait interval a cheap delta resync; a replica that returns after it gets a full rebuild. The down-clock starts when the volume detaches (just before reboot), so a normal reboot easily fits the window — but a node left down for a long maintenance action will pay the full-rebuild cost on every replica it hosted. (A few volume types disable the revision counter for write throughput and always full-rebuild regardless — harmless, just slower.)

Extra pre-reboot gates when the node hosts a database

Beyond "storage healthy, cluster green," clear these before rebooting a node that hosts a clustered database:

  • Galera (MariaDB): confirm the cluster is fully synced (all members Synced, full membership) and that a state transfer could actually run, before the reboot — and wait for it to re-sync fully after, before moving to the next node. A Galera StatefulSet showing replicas: 0 is usually the operator mid-recovery (it scales down to pick a bootstrap node by sequence number, then scales back), not a failure — leave it alone.
  • Primary/replica databases (e.g. Postgres): reboot a node holding only a replica freely; reboot the node holding the primary last, so automatic failover relocates it once, cleanly.

After the upgrade

  • Bump the local tooling to match the cluster: talosctl, and longhornctl after a Longhorn bump.
  • Update the version references in the repo.

See also: Storage · Operations · Architecture · Storage recovery · Database operations & recovery.