How to Choose Kubernetes Storage for High IOPS Workloads
April Wong
The best storage for stateful workloads on Kubernetes meets the application's sustained I/O and recovery requirements with an operating model your team can support. A managed cloud volume is a useful starting point when you want the provider to run the storage service. OpenEBS, Rook/Ceph and Portworx provide different ways to manage storage through Kubernetes. Nirvana combines NKS with ABS-backed persistent storage for teams evaluating its cloud as a complete stack.
A CSI driver name or peak IOPS number is not enough to choose between them. Compare the underlying storage, replication policy, node limits and recovery path.
Understand what a persistent volume promises
In Kubernetes, a PersistentVolume is a storage resource whose lifecycle is separate from an individual Pod. A PersistentVolumeClaim requests storage for an application. A StorageClass defines a provisioning class; its implementation determines what storage is created. The Container Storage Interface, or CSI, connects Kubernetes to a storage system. These abstractions organize access. They do not make every backend equally fast or durable. See the Kubernetes persistent-volume documentation.
That distinction matters during failure. Retaining a volume can preserve data after a Pod stops, while a replacement Pod still has to be scheduled, gain access to storage and recover the application. A running Pod is not proof that a database has completed recovery.
Compare the storage systems at the right layer
| Option | What you are choosing | Main evaluation question |
|---|---|---|
| Cloud block volumes through CSI | A provider-operated storage service attached to cluster nodes | Do volume and instance limits meet the workload in the required region? |
| Nirvana NKS with ABS | Managed Kubernetes on Nirvana with its integrated block-storage layer | Does the complete stack meet sustained performance, placement and recovery requirements? |
| OpenEBS Local PV | Kubernetes-managed access to local storage | Does the application supply the replication and recovery required when a node is unavailable? |
| OpenEBS Replicated Storage | A replicated storage engine managed through Kubernetes | What are the latency, capacity and rebuild costs at the selected replica count? |
| Rook/Ceph | An operator that deploys and manages Ceph | Can the team operate the Ceph topology, network, capacity and recovery processes? |
| Portworx Enterprise | A Kubernetes storage platform with configurable volume and replication policies | Which deployment, protection and support features are required, and what do they cost? |
This is an architecture comparison, not a performance league table. Sources: NKS, OpenEBS architecture, OpenEBS replication, Rook/Ceph and Portworx operations.
Local storage can be a sensible choice when the application already replicates data across independent nodes. Additional storage replication may improve a different failure boundary, but it also consumes capacity and resources. Draw the failure domains before deciding how many copies to keep at each layer.
Start with the workload rather than the volume label
Measure the I/O that reaches the disk. A database writing small synchronous transactions stresses a different path from an analytical scan reading large blocks. Record read/write mix, block sizes, concurrency, working-set size and required durability.
Then define the service objective. For a transactional application, tail latency may matter more than maximum throughput. For an indexer, the constraint may be keeping ingestion current while reads continue. An agent checkpoint store may need both durable commit latency and quick worker replacement.
Ask whether the advertised performance is a baseline, a configured limit or temporary headroom. Record the node's storage and network limits too. A volume capable of more I/O cannot exceed every constraint elsewhere in the path.
Run a useful persistent-volume comparison
Use a disposable test volume and publish the configuration with the result. A reproducible evaluation records Kubernetes and CSI versions, node shape, volume size, filesystem, replica count, encryption settings, region and relevant network topology. Keep these matched where possible; explain differences where architectures require them.
Test small random reads, small random writes and a mixed profile that resembles the application. Add a separate sequential test when scans or large transfers matter. For durability-sensitive work, measure the flush or commit barrier explicitly. Direct I/O and durable application acknowledgement are related but different properties.
For each profile, report sustained IOPS, throughput and p50/p95/p99 latency across a long enough interval to reveal queue growth, background work and any documented burst behavior. Show latency alongside throughput: raising queue depth may increase completed I/O while making individual requests slower.
Follow the synthetic test with the real application. Keep the same transaction, query or checkpoint correctness conditions. The useful result is whether the deployment meets the service objective under representative concurrency.
Include recovery in the acceptance criteria
A production evaluation should include a clean Pod restart, node replacement, storage reattachment and application recovery. Run these in an isolated test environment with disposable data or validated backups. Test backup restoration separately from worker replacement.
Record the last acknowledged write before interruption and verify it afterward. Measure from failure detection to the first correct application response, then to steady performance. Include the time needed for caches and replicas to recover.
Check access modes, topology constraints, reclaim policy and volume expansion before migration. A claim configured for one writer is not a shared filesystem simply because multiple Pods can reference it. Retaining storage also does not protect against every accidental deletion or regional failure.
Where NKS and ABS belong in the shortlist
Nirvana Kubernetes Service integrates managed Kubernetes with Nirvana compute and ABS persistent volumes. The docs identify Silicon Valley, us-sva-2, as the current NKS region. ABS documentation lists 20,000 baseline IOPS. These are product specifications to validate against the required deployment, not results from a matched test against every storage system above.
Consider the stack for storage-intensive applications when that region and operating model fit. Compare the complete deployment bill, including compute, retained capacity, backups and networking. Do not apply a storage-only price to the entire cluster.
Nirvana's existing storage and agent evidence can inform a proof of concept, but it does not establish a universal CSI ranking. A claim that NKS outperforms Portworx, OpenEBS or Rook/Ceph requires equivalent, reproducible tests of those configurations.
Common questions
Which CSI driver is fastest? There is no useful universal answer without naming the backend and workload. Test the whole I/O path at the required durability and replica settings.
Does a StatefulSet back up the data? No. Stable identity and storage association do not replace a backup and restore process.
Do I need 20,000 IOPS? Only if your measured demand and headroom justify it. Also check throughput and latency; a scan can need bandwidth without needing the same small-I/O rate as a transaction workload.
How do I evaluate Nirvana? Start with the NKS documentation and ABS specifications. Bring a workload profile, recovery objective and deployment region so the proof of concept answers a real production question.
Related Posts

How to Choose Cloud Infrastructure for a Vector Database
There is no universal IOPS requirement for vector search. This guide shows how to size the engine, find the bottleneck and compare hosting at a fixed recall and latency target.

Which Cloud Is Best for Long-Running AI Agents?
Choose agent hosting by the work it preserves. Compare persistence, recovery, operating models and cost, with clearly scoped Nirvana benchmark evidence.

What Is a Harness? The Agent, The Model, the Harness, and the Sandbox
An agent harness turns a model's decisions into action: it runs the loop, executes each tool call, keeps the notes and calls time. The word is new, the thing is not. What changed in 2026 is that OpenAI, Vercel and LangChain started selling it.