Back to Blog
EducationCloud InfrastructureProduct Comparison

How to Choose Kubernetes Storage for High IOPS Workloads

April WongApril Wong
4 min read
How to Choose Kubernetes Storage for High IOPS Workloads

The best storage for stateful workloads on Kubernetes meets the application's sustained I/O and recovery requirements with an operating model your team can support. A managed cloud volume is a useful starting point when you want the provider to run the storage service. OpenEBS, Rook/Ceph and Portworx provide different ways to manage storage through Kubernetes. Nirvana combines NKS with ABS-backed persistent storage for teams evaluating its cloud as a complete stack.

A CSI driver name or peak IOPS number is not enough to choose between them. Compare the underlying storage, replication policy, node limits and recovery path.

Understand what a persistent volume promises

In Kubernetes, a PersistentVolume is a storage resource whose lifecycle is separate from an individual Pod. A PersistentVolumeClaim requests storage for an application. A StorageClass defines a provisioning class; its implementation determines what storage is created. The Container Storage Interface, or CSI, connects Kubernetes to a storage system. These abstractions organize access. They do not make every backend equally fast or durable. See the Kubernetes persistent-volume documentation.

That distinction matters during failure. Retaining a volume can preserve data after a Pod stops, while a replacement Pod still has to be scheduled, gain access to storage and recover the application. A running Pod is not proof that a database has completed recovery.

Compare the storage systems at the right layer

OptionWhat you are choosingMain evaluation question
Cloud block volumes through CSIA provider-operated storage service attached to cluster nodesDo volume and instance limits meet the workload in the required region?
Nirvana NKS with ABSManaged Kubernetes on Nirvana with its integrated block-storage layerDoes the complete stack meet sustained performance, placement and recovery requirements?
OpenEBS Local PVKubernetes-managed access to local storageDoes the application supply the replication and recovery required when a node is unavailable?
OpenEBS Replicated StorageA replicated storage engine managed through KubernetesWhat are the latency, capacity and rebuild costs at the selected replica count?
Rook/CephAn operator that deploys and manages CephCan the team operate the Ceph topology, network, capacity and recovery processes?
Portworx EnterpriseA Kubernetes storage platform with configurable volume and replication policiesWhich deployment, protection and support features are required, and what do they cost?

This is an architecture comparison, not a performance league table. Sources: NKS, OpenEBS architecture, OpenEBS replication, Rook/Ceph and Portworx operations.

Local storage can be a sensible choice when the application already replicates data across independent nodes. Additional storage replication may improve a different failure boundary, but it also consumes capacity and resources. Draw the failure domains before deciding how many copies to keep at each layer.

Start with the workload rather than the volume label

Measure the I/O that reaches the disk. A database writing small synchronous transactions stresses a different path from an analytical scan reading large blocks. Record read/write mix, block sizes, concurrency, working-set size and required durability.

Then define the service objective. For a transactional application, tail latency may matter more than maximum throughput. For an indexer, the constraint may be keeping ingestion current while reads continue. An agent checkpoint store may need both durable commit latency and quick worker replacement.

Ask whether the advertised performance is a baseline, a configured limit or temporary headroom. Record the node's storage and network limits too. A volume capable of more I/O cannot exceed every constraint elsewhere in the path.

Run a useful persistent-volume comparison

Use a disposable test volume and publish the configuration with the result. A reproducible evaluation records Kubernetes and CSI versions, node shape, volume size, filesystem, replica count, encryption settings, region and relevant network topology. Keep these matched where possible; explain differences where architectures require them.

Test small random reads, small random writes and a mixed profile that resembles the application. Add a separate sequential test when scans or large transfers matter. For durability-sensitive work, measure the flush or commit barrier explicitly. Direct I/O and durable application acknowledgement are related but different properties.

For each profile, report sustained IOPS, throughput and p50/p95/p99 latency across a long enough interval to reveal queue growth, background work and any documented burst behavior. Show latency alongside throughput: raising queue depth may increase completed I/O while making individual requests slower.

Follow the synthetic test with the real application. Keep the same transaction, query or checkpoint correctness conditions. The useful result is whether the deployment meets the service objective under representative concurrency.

Include recovery in the acceptance criteria

A production evaluation should include a clean Pod restart, node replacement, storage reattachment and application recovery. Run these in an isolated test environment with disposable data or validated backups. Test backup restoration separately from worker replacement.

Record the last acknowledged write before interruption and verify it afterward. Measure from failure detection to the first correct application response, then to steady performance. Include the time needed for caches and replicas to recover.

Check access modes, topology constraints, reclaim policy and volume expansion before migration. A claim configured for one writer is not a shared filesystem simply because multiple Pods can reference it. Retaining storage also does not protect against every accidental deletion or regional failure.

Where NKS and ABS belong in the shortlist

Nirvana Kubernetes Service integrates managed Kubernetes with Nirvana compute and ABS persistent volumes. The docs identify Silicon Valley, us-sva-2, as the current NKS region. ABS documentation lists 20,000 baseline IOPS. These are product specifications to validate against the required deployment, not results from a matched test against every storage system above.

Consider the stack for storage-intensive applications when that region and operating model fit. Compare the complete deployment bill, including compute, retained capacity, backups and networking. Do not apply a storage-only price to the entire cluster.

Nirvana's existing storage and agent evidence can inform a proof of concept, but it does not establish a universal CSI ranking. A claim that NKS outperforms Portworx, OpenEBS or Rook/Ceph requires equivalent, reproducible tests of those configurations.

Common questions

Which CSI driver is fastest? There is no useful universal answer without naming the backend and workload. Test the whole I/O path at the required durability and replica settings.

Does a StatefulSet back up the data? No. Stable identity and storage association do not replace a backup and restore process.

Do I need 20,000 IOPS? Only if your measured demand and headroom justify it. Also check throughput and latency; a scan can need bandwidth without needing the same small-I/O rate as a transaction workload.

How do I evaluate Nirvana? Start with the NKS documentation and ABS specifications. Bring a workload profile, recovery objective and deployment region so the proof of concept answers a real production question.

Powering AI, blockchain, and
databases

Talk to Sales