Skip to content

Runbooks

Generic, sanitized operational runbooks. Deliberately specifics-free — no hostnames, IPs, or secret material; exact commands live in the relevant repo's README.

Available

  • Node upgrade & maintenance — Talos/Kubernetes upgrade discipline, the storage/health preflight, and safe worker vs. control-plane reboots.
  • Storage recovery — stale VolumeAttachments after a reboot, salvaging from a surviving replica without silent data loss, and when to stop patching.
  • Mail deliverability — proving DKIM actually signs, selector rotation, and shared-secret (TURN) convergence.
  • Database operations & recovery — Galera and operator-managed database recovery hazards, and the cert-manager/operator Secret-ownership trap.
  • Capacity & resource limits — reading OOM/probe-storm signals on small nodes, and why a memory limit is not a reservation.
  • Self-managed Argo CD upgrades — upgrading the GitOps controller that manages itself, without a cascade.