Skip to content

Database operations & recovery

Durable operating hazards for the clustered databases. The exact recovery commands live in your operations runbooks; this page is the judgement that keeps a recovery from making things worse.

Galera: never start a node without its replication provider

Starting mariadbd with the Galera provider disabled destroys recently-replicated data

Starting the server with --wsrep-provider=none against a live Galera data directory runs plain InnoDB recovery, which rolls back the XA-prepared transactions — precisely the most recently replicated rows. This permanently destroys data Galera had applied. Never do it for "debugging." If you genuinely need emergency access to a Galera volume, --skip-grant-tables (which preserves XA state) or mounting the PVC read-only in a debug pod are the safe paths — never combined with disabling the provider.

Galera operator recovery: verify the bootstrap node before unsuspending

When a Galera operator recovers a cluster, the bootstrap node must be the one with the highest sequence number and a non-zeroed cluster UUID. If it bootstraps from a zeroed/empty node, the healthy nodes state-transfer from the empty one — copying emptiness over good data. Check the operator's recovery selection before letting it proceed. And remember a Galera StatefulSet at replicas: 0 is usually the operator mid-recovery (scaling down to read each node's saved sequence number, then back up), not a failure — leave it alone.

Two controllers must never own the same Secret

An operator that rewrites a cert-manager-managed Secret causes a re-issue loop

If an operator does server-side-apply on a TLS Secret that cert-manager also manages, it can strip cert-manager's revision-tracking annotation on every reconcile; cert-manager then treats the Secret as stale and re-issues — a tight loop (many issues per second) that can cascade into API-server load and dependent-controller failures. This is a general "two controllers fighting over one object" hazard, not a single-product bug. Where an operator offers an externally-managed / custom TLS Secret mode (it reads the Secret instead of reconciling it), prefer that — at the cost of losing cert hot-reload, so pair a cert renewal with a rolling restart.

Cross-version operator upgrades surface latent state

A long-lived database cluster accumulates state (schema objects, functions, stored revisions) from every operator version it has ever run. Operators often hash-skip work they believe is already applied — which can hide a chronically-failing step for a very long time. A major-version operator upgrade that re-evaluates that skipped state can suddenly surface a latent problem as a reconcile loop "that started at the upgrade" but was actually dormant for months. When debugging one, check whether the step was ever succeeding before the upgrade, and — because operators frequently discard the sub-process stderr from their exec calls — reproduce the failing command manually with the same arguments to see the real error. Take a fresh backup through the operator's own tested backup path before any major upgrade.

See also: Node upgrade & maintenance · Storage recovery · Renovate.