A stateful Kubernetes service is recoverable when it can reopen the right data and serve correct requests after a failure. A replacement pod reaching Running is only part of that process. Storage attachment, database recovery, application readiness and client reconnection can each extend the outage.
TL;DR
- A replacement pod reaching Running does not prove that the application recovered its data correctly.
- Test process restart, pod replacement, planned maintenance and abrupt node loss as separate recovery paths.
- Verify snapshot and expansion support in the installed driver and StorageClass; a persistent volume is not a backup.
- Rehearse an independent restore, check acknowledged writes and measure recovery against the application's data-loss and downtime limits.
For NKS, turn production readiness into three demonstrated outcomes: the workload can recover after compute disruption, an independent recovery copy can be restored, and capacity can grow through the installed storage stack. The tests below separate those outcomes so a successful restart is not mistaken for a complete recovery plan.
Updated 7 October 2026. This is an operator runbook, not a report of tests executed on a customer cluster. Run disruptive exercises only in an isolated test environment with disposable data.
What does NKS document, and what must the cluster prove?
Nirvana's NKS storage guide documents ABS-backed persistent volumes in us-sva-2 and a preconfigured StorageClass for PVC provisioning. It does not, by itself, establish a tested backup policy, restore time or expansion workflow for every installed cluster configuration.
Begin with a read-only inventory. The following standard Kubernetes commands inspect the current cluster; resource names and permissions depend on your environment:
kubectl get storageclass -o yaml
kubectl get csidriver
kubectl get pvc -A
kubectl get pv
kubectl api-resources --api-group=snapshot.storage.k8s.io
kubectl get volumesnapshotclass
If a snapshot resource is unavailable or access is denied, record that result rather than assuming support. Confirm the driver and version with the platform owner. Keep the inventory with the application's deployment revision so a later upgrade does not silently invalidate the runbook.
Inspect the actual PV as well as its StorageClass. Kubernetes documents that dynamically provisioned volumes inherit a reclaim policy, and Delete can remove the backing storage. Also, ReadWriteOnce means one node, not necessarily one pod. Those details affect deletion procedures and concurrent writers.
Can the workload recover after losing its compute?
Create a representative test dataset and a small stream of uniquely numbered committed writes. Keep an independent record of the writes the client received confirmation for. Before the disruption, record the pod identity, node, PVC/PV binding and latest confirmed sequence number.
Run these as distinct exercises, each with a defined expected result:
| Exercise | What to observe | What a pass establishes |
|---|---|---|
| Application process restart | Startup logs and recovered transaction state | The application can reopen its existing data |
| Pod replacement | New pod identity, mounted claim and correct reads/writes | The controller and volume path support replacement |
| Planned node maintenance | Eviction behavior, replacement placement and service continuity | The maintenance procedure works for this workload |
| Abrupt worker loss | Detection, rescheduling, attachment and database recovery | The tested failure path meets the recovery target |
Do not force-delete a stateful pod or bypass attachment protections simply to make a test finish. Resolve the previous writer's status and the storage driver's documented procedure; two writers that both believe they own a database can be more damaging than a longer outage.
StatefulSets provide stable workload identity and storage relationships. They do not turn a database into a replicated service. Configure and test the database's own replication and failover behavior if the availability requirement needs it.
Likewise, a PodDisruptionBudget is relevant to voluntary disruptions; it does not prevent a machine from failing. A successful drain and a successful abrupt-failure test should remain separate entries in the evidence.
Measure until a client can perform a correct state-dependent read and a new committed write. Record the phase that consumed the most time: failure detection, scheduling, mount, database replay or readiness. This identifies a fix more clearly than one undifferentiated “restart time.”
Can you restore data after deletion or corruption?
Retention of a live volume is not historical recovery. A mistaken application delete can be written successfully to perfectly durable storage. Define the recovery point objective (how much recent data may be lost) and recovery time objective (how long restoration may take) for the application.
Kubernetes volume snapshots require snapshot APIs, a controller and support from the CSI driver. The existence of a PVC is not proof that those components are installed. Snapshot topology and restore destination also need verification.
Choose a database-appropriate consistency method. For PostgreSQL, its backup documentation distinguishes SQL dumps, filesystem-level backups and continuous archiving. Select and rehearse a supported method for the recovery objective; a copied directory from a running database is not automatically a valid backup.
A useful restore rehearsal starts from a fresh target, not the original live volume. Restore the data, supply the required configuration and keys, and validate application-level results. Compare known record counts, recent sequence numbers and representative queries. Prove that the restore can proceed using the documented credentials and recovery materials without relying on the failed worker.
Record backup age, copy location, restored point and elapsed time. Test the procedure after meaningful version or schema changes. A dashboard showing “backup succeeded” is evidence of a backup operation, not evidence that the application can recover.
Can the volume expand before the workload runs out of space?
Kubernetes expansion depends on the StorageClass's allowVolumeExpansion setting and driver support. The documented path is to increase the PVC's requested capacity; the backing volume is expanded rather than replaced. Confirm filesystem and driver behavior for the deployed versions before scheduling the exercise.
In a disposable environment, grow a volume under a representative write load. Observe the PVC conditions, events, reported capacity and the filesystem capacity visible inside the application. These are separate checks: a larger control-plane value does not alone demonstrate usable free space.
Continue the workload after growth and confirm data integrity. Record any restart requirement and include it in the maintenance plan. Do not treat shrinking as the rollback mechanism; retain an independent recovery copy and a tested restore route.
Capacity planning needs working space, not only stored rows. Track the workload's temporary files, logs and compaction or merge behavior. Set an alert early enough to leave time for the actual expansion procedure, including any approval or support dependency.
What should the production recovery record contain?
Keep one concise record per workload: storage class and driver, image and database versions, replica layout, disruption procedure, backup policy, restore instructions, expansion evidence and named operational owners. Add the dates and measured outcomes of the last rehearsal.
The acceptance gate is specific: confirmed data survived the tested compute failure; a separate recovery copy restored correctly within the agreed objectives; growth became usable without corrupting the application. Any untested failure domain remains an explicit open item.
Use the NKS storage guide to establish the provisioning path, then review these operational requirements with the Nirvana team. Infrastructure performance and demonstrated recovery solve different parts of production readiness.
FAQ
Does a StatefulSet include database backup or replication?
No. Those need an application-specific design and operating procedure.
Does deleting a pod delete its PVC?
They are separate resources, but controllers and retention policies matter. Inspect the workload and storage policies before testing any deletion path.
Is a snapshot enough to claim disaster recovery?
Only after verifying consistency, retention, failure isolation, required keys and a successful restore in the intended recovery location.
About Nirvana Labs
Nirvana Labs is a high-performance storage cloud purpose built for blockchain, AI and databases i.e. the most demanding, real-time, stateful workloads. Accelerated Block Storage (ABS) offers 20K baseline IOPS included, no over provisioning. Nirvana Kubernetes Service (NKS) with Karpenter auto-scaling, high clock-speed compute and private networking. Backed by Jump Trading, Crucible, etc with 50+ customers live in production today.
Learn more at Nirvana Labs
Nirvana Cloud | Pricing | Blog | Docs | Changelog | LinkedIn | Twitter | Telegram | YouTube
Related Posts

How to Benchmark PostgreSQL WAL and Commit Performance on Cloud Storage
A repeatable PostgreSQL storage evaluation: control durability settings, separate transaction and commit latency, inspect WAL and checkpoint metrics, and test recovery.

How to Test Redis AOF Rewrites and Recovery Before Production
A practical acceptance test for Redis persistence: measure rewrites under load, validate acknowledged writes after failure, and prove independent backup recovery.

Altinity.Cloud BYOC: An Evaluation and Migration Guide
Evaluate Altinity.Cloud BYOC for ClickHouse® with ownership, workload, backup and migration checks before a reversible cutover.

