INFRA Signal 88
Kubernetes disaster recovery: Guidance from three reproducible failure scenarios
CNCF guidance demonstrates how backups alone do not guarantee recoverability in Kubernetes stateful applications through three lab-tested failure scenarios
Disaster recovery for Kubernetes stateful workloads often fails at the boundaries between layers, backups, cluster infrastructure, application definitions, and data, rather than within them. This guidance provides concrete, reproducible scenarios to test recovery procedures beyond backup completion, exposing gaps that teams typically discover only during real outages. The material is tool-agnostic but grounded in real implementations, making it actionable for operators regardless of their stack.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Backup completion does not guarantee data recoverability; verification must confirm volume bytes were actually moved to external storage
GitOps controllers can restore application definitions without data, leaving databases running but empty after recovery
Disaster recovery plans must explicitly define what infrastructure backups are restored into, as backup tools do not provision clusters or networks
THE READ
What the cluster adds up to.
The guidance distinguishes between having backups and being able to recover by isolating three failure modes that occur at the joins between recovery layers. Each scenario is designed to be reproducible on a laptop using a lab environment with two local Kubernetes clusters, an external S3-compatible store, and a Git service. The lab runs a PostgreSQL workload with known data, allowing recovery outcomes to be validated against expected results rather than dashboard indicators. This approach shifts disaster recovery testing from theoretical checklists to concrete, observable failures.
The first scenario demonstrates that backup tools report completion without confirming that volume data actually moved to external storage. The lab shows a Velero data mover reporting the exact byte count transferred, a metric that backup tools should expose but often do not. Even when data is confirmed to have moved, recovery still depends on application consistency hooks, storage class mappings, and infrastructure compatibility. The guidance emphasizes that a backup’s 'Completed' status is not evidence of recoverability, only an end-to-end recovery test can provide that.
The second scenario exposes a critical gap in GitOps-driven recovery: controllers can restore application definitions without restoring the corresponding data. In the lab, syncing a PostgreSQL StatefulSet from Git results in a running but empty database, as the Git repository holds only the declared state, not the stored state. This failure mode is invisible to dashboards and requires explicit validation of data integrity. The guidance recommends treating GitOps and backup tools as complementary, not interchangeable, and testing their interaction under failure conditions.
The third scenario, though truncated in the material, highlights the boundary between backup tools and infrastructure recovery. Backup tools restore resources into an existing cluster but do not provision the cluster itself, its nodes, or its networking. Disaster recovery plans must explicitly define what infrastructure backups are restored into, whether through infrastructure-as-code or Cluster API. The guidance frames this as a responsibility split: backup tools handle application and data recovery, while other tooling must handle the underlying Kubernetes infrastructure. Teams must test both in concert to avoid silent failures.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗