Skip to content

Upgrade methodology

Upgrades on a small cluster are where things break if you improvise. alpstack follows a few fixed rules so every upgrade is boring.

OS and Kubernetes move together, one minor at a time

Talos ships a matched Kubernetes version range: a given Talos release supports a bounded set of Kubernetes minors. Treat them as coupled — plan the Talos upgrade and the Kubernetes upgrade as one exercise, and never skip a minor in either. Skipping puts you outside the tested range and turns a routine step into a debugging session.

The practical order for a minor bump:

  1. Clear the prerequisites first (see gating below).
  2. Upgrade Talos on the nodes (control plane, then workers), staying within its supported Kubernetes range.
  3. Bump the Kubernetes control-plane version, let it settle, then the workers.
  4. Verify the whole stack is healthy before starting the next minor.

Gate every upgrade on its prerequisites

The failure mode to avoid is bumping the core and discovering an add-on can't run on it. So the core version is gated — it doesn't move until its dependencies are ready:

  • The CNI must support the target Kubernetes version before the control plane is bumped.
  • Operators and CRD-based add-ons (storage, databases, ingress) must have a compatible release available.
  • Anything with a known upstream blocker holds the whole chain until the blocker clears.

Write the gate down as an explicit checklist per upgrade — "core X.Y is blocked on CNI ≥ A.B and operator ≥ C.D" — so the dependency order is visible and nobody bumps the core early.

Renovate does the reconnaissance

Automated dependency PRs surface new component versions as they appear, so the gating checklist is mostly a matter of reading the queue rather than hunting for releases. See Renovate.

Draining nodes with node-local volumes

Most volumes here are backed by the node's local disk. A node reboot or drain therefore has to respect the data, not just the pods:

  • One node at a time. Cordon, drain, upgrade/reboot, and confirm the node and its volumes are healthy again before touching the next one. Never drain two nodes carrying replicas of the same volume together.
  • Respect PodDisruptionBudgets and volume health. For a replicated volume, make sure a healthy replica exists elsewhere before you evict the pod using it. For a genuinely single-replica volume, accept that its workload is down for the duration and schedule accordingly.
  • Detach cleanly. Let the CSI layer detach a volume before the node goes down, so you don't chase stale attachments on the way back up. See Storage recovery.

Control plane vs. workers

Control-plane nodes run etcd — keep quorum. With three control-plane nodes you can take one at a time; never two. Workers have no such constraint beyond volume health, but the one-at-a-time discipline still applies so capacity and replicas stay intact.

Self-managed platform components

Some platform pieces reconcile themselves — most notably the GitOps controller. Upgrading those has an extra ordering wrinkle; see the runbook Self-managed Argo CD upgrades.

See also: Node upgrade & maintenance · Component overview.