Database operations & recovery¶
Durable operating hazards for the clustered databases. The exact recovery commands live in your operations runbooks; this page is the judgement that keeps a recovery from making things worse.
Galera: never start a node without its replication provider¶
Starting mariadbd with the Galera provider disabled destroys recently-replicated data
Starting the server with --wsrep-provider=none against a live Galera data directory
runs plain InnoDB recovery, which rolls back the XA-prepared transactions — precisely
the most recently replicated rows. This permanently destroys data Galera had applied.
Never do it for "debugging." If you genuinely need emergency access to a Galera volume,
--skip-grant-tables (which preserves XA state) or mounting the PVC read-only in a debug
pod are the safe paths — never combined with disabling the provider.
Galera operator recovery: verify the bootstrap node before unsuspending¶
When a Galera operator recovers a cluster, the bootstrap node must be the one with the
highest sequence number and a non-zeroed cluster UUID. If it bootstraps from a zeroed/empty
node, the healthy nodes state-transfer from the empty one — copying emptiness over good
data. Check the operator's recovery selection before letting it proceed. And remember a Galera
StatefulSet at replicas: 0 is usually the operator mid-recovery (scaling down to read
each node's saved sequence number, then back up), not a failure — leave it alone.
Two controllers must never own the same Secret¶
An operator that rewrites a cert-manager-managed Secret causes a re-issue loop
If an operator does server-side-apply on a TLS Secret that cert-manager also manages, it can strip cert-manager's revision-tracking annotation on every reconcile; cert-manager then treats the Secret as stale and re-issues — a tight loop (many issues per second) that can cascade into API-server load and dependent-controller failures. This is a general "two controllers fighting over one object" hazard, not a single-product bug. Where an operator offers an externally-managed / custom TLS Secret mode (it reads the Secret instead of reconciling it), prefer that — at the cost of losing cert hot-reload, so pair a cert renewal with a rolling restart.
Cross-version operator upgrades surface latent state¶
A long-lived database cluster accumulates state (schema objects, functions, stored revisions) from every operator version it has ever run. Operators often hash-skip work they believe is already applied — which can hide a chronically-failing step for a very long time. A major-version operator upgrade that re-evaluates that skipped state can suddenly surface a latent problem as a reconcile loop "that started at the upgrade" but was actually dormant for months. When debugging one, check whether the step was ever succeeding before the upgrade, and — because operators frequently discard the sub-process stderr from their exec calls — reproduce the failing command manually with the same arguments to see the real error. Take a fresh backup through the operator's own tested backup path before any major upgrade.
See also: Node upgrade & maintenance · Storage recovery · Renovate.