How to Choose Cloud Infrastructure for a Vector Database
April Wong
Choose cloud infrastructure for a vector database by matching the engine's memory model, search-quality target, query concurrency and recovery requirements. There is no universal IOPS requirement for vector search. A mostly memory-resident index may be CPU- or memory-bound; disk-backed vectors, filtered retrieval, ingestion and index rebuilding can place substantial demand on storage.
The first useful question is where time goes in your own retrieval pipeline. A faster disk cannot remove latency spent creating an embedding, reranking results or waiting for an external model.
A limitation to know before choosing Nirvana: Our LangChain benchmark found slower per-query Qdrant tail latency on Nirvana ABS than on AWS io2, even though Nirvana completed the combined agent workload sooner. A high IOPS figure or faster overall agent run is not evidence of the fastest vector search. If search latency is your main requirement, evaluate that result separately.
Separate database choice from hosting choice
Qdrant, Weaviate, Milvus and pgvector make different architectural choices. You can compare their search capabilities and operating model before deciding where to run them.
| Engine | Infrastructure considerations | A useful starting point |
|---|---|---|
| Qdrant | Memory placement for vectors, indexes and payloads changes the RAM and disk tradeoff | Measure the selected memory tier and quantization settings on your dataset |
| Weaviate | Its HNSW index requires memory; compression and the selected index type affect resource use | Size the index and concurrent workload before assuming more disk performance is the answer |
| Milvus | A distributed architecture separates compute and storage responsibilities | Include the deployed version's full component and storage requirements in the plan |
| PostgreSQL with pgvector | Vector search shares a database with relational data, writes, maintenance and SQL queries | Check index behavior and contention alongside normal PostgreSQL operations |
Primary references: Qdrant memory tiers, Weaviate resource planning, Milvus architecture and pgvector documentation.
A managed database transfers much of the operational work to the service provider. Self-hosting gives you more control over placement, configuration and infrastructure, while leaving upgrades, replication, monitoring and recovery with your team. A cloud VM is not automatically a managed vector database.
Size the data before sizing the machine
As an illustrative lower bound, ten million vectors with 768 dimensions stored as float32 values contain 30.72 GB of raw vector values: 10,000,000 × 768 × 4 bytes. That is about 28.6 GiB. It excludes index structures, payloads, metadata, replication, logs, temporary build space and operating-system overhead.
Do not turn that arithmetic into a RAM recommendation. The index type, engine implementation, compression and cache behavior determine what must remain in memory. Leave room for ingestion and rebuilds as well as steady-state queries.
Qdrant's storage documentation describes configurable memory placement. Weaviate's guidance describes its HNSW memory requirements. Those are reasons to test engine-specific settings, rather than using one capacity multiplier for every database.
Find the bottleneck in the retrieval pipeline
Break an end-to-end request into embedding generation, network transit, database search, filtering or payload reads, reranking and answer generation. Measure each stage separately and under realistic concurrency.
If database search slows while CPU is saturated, examine query and index settings before buying more IOPS. If memory pressure causes frequent disk access and storage queues grow, a faster storage path or a different memory layout may help. If search is fast but the full request is slow, focus on the stage actually consuming time.
Keep the required search quality fixed. Recall measures how many relevant nearest neighbours the approximate search finds compared with the chosen reference. A faster result at substantially lower recall is not an equivalent result. Include real filters and tenant distributions; an unfiltered demonstration can miss the expensive production path.
Use IOPS as a measured requirement
Observe storage reads and writes while the database serves the target query rate and ingestion load. Record block sizes, cache misses, queue depth, latency and CPU utilization. Repeat after a restart and with a working set larger than effective memory.
You can use a simple planning identity: storage operations per second are approximately queries per second multiplied by the average storage operations required per query, plus writes and background work. The operations-per-query term must come from measurement. It changes with caching, filters, index layout and data distribution.
This explains why neither vector count nor a provider's peak IOPS number can determine application performance on its own. The meaningful target is the required query rate and recall at the agreed p95/p99 latency.
Why queue depth matters for vector search
Queue depth is the number of outstanding storage operations at the layer being measured. At QD=1, that I/O stream waits for one operation to finish before issuing another. Its performance depends heavily on individual-operation latency. A high-queue-depth throughput test measures a different situation, with many requests in flight.
The LangChain report describes its per-query Qdrant tradeoff as QD=1. Its raw 4 KB random-read fio test uses QD=256. Those numbers cannot be substituted for each other. A search that follows dependent reads through an HNSW graph may have limited I/O parallelism along that path.
Batching and concurrent queries can create more parallel work, but do not automatically shorten a single chain of dependent reads. Many agents can also create aggregate device concurrency while individual query paths remain latency-sensitive. Measure the queue depth actually reached; do not assume every vector engine or query always runs at QD=1.
Compare infrastructure with one test plan
Use the same embedding set, distance metric, search-quality target, filters and query distribution on each option. Pin the engine version and index settings. Record both warm behavior and a defined cold or recovery state.
Test single-query latency and throughput at realistic concurrency separately. Record actual storage queue depth. Measure query throughput, p50/p95/p99 latency, recall, ingestion rate, time until new data is searchable, index-build time and storage consumed. Run queries while ingestion and maintenance continue. Report failures and timeouts alongside completed requests.
Then replace a worker or restore from backup in a test environment. Confirm the dataset and index are correct and record how long it takes to return to the service objective. Include replica placement and the failure boundary being tested; a successful local restart does not prove multi-region disaster recovery.
Price the smallest configuration that passes. Include replicas, memory, compute, block and object storage where required, snapshots, backups, network transfer and operational effort. Useful comparisons are cost at a fixed dataset and query rate, or cost per successful query at a fixed recall and latency target. A bare monthly VM price leaves too much out.
What Nirvana's LangChain benchmark actually shows
The published LangChain read benchmark compares a mixed Qdrant, Redis and PostgreSQL workload across Nirvana ABS and four AWS storage configurations. The R3 production-scale run uses five million 768-dimensional vectors across 50 collections, Qdrant 1.16 with inline storage and INT8 scalar quantization, and 1,000 agents completing 100 tasks each. Each task includes two operations per service.
| Metric at 100,000 tasks, R3 | Nirvana ABS | AWS io2-64k |
|---|---|---|
| Qdrant vector-search p99 | 301 ms | 140 ms |
| Complete agent-task p99 | 725 ms | 771 ms |
| Total workload completion | 58 minutes | 69 minutes |
Lower is better in every row. Nirvana's Qdrant p99 was about 2.15 times the io2 result in this run. The complete workload nevertheless finished about 16% sooner. These are different outcomes: a throughput advantage across the mixed workload does not establish a per-query vector-search advantage. Per-service p99 values also cannot simply be added to reconstruct task p99.
The configurations had the same nominal vCPU count, RAM capacity and volume size, but used different compute platforms. This is a comparison of the tested deployments, not an experiment isolating storage batching as the sole cause. It does not establish the same result for Weaviate, Milvus, pgvector or another Qdrant configuration.
When to evaluate Nirvana
For an application whose main requirement is low tail latency on individual vector searches, the published result favors the tested io2 configuration. Start with query latency at your required recall and concurrency; do not choose ABS on peak IOPS alone.
For mixed agent workloads where overall completion time and sustained throughput matter, the LangChain result provides a reason to evaluate Nirvana. Confirm that the vector-search component still meets your application's latency objective. Ingestion, index builds and recovery need their own measurements; this result does not establish a Nirvana advantage for them.
ABS and NKS are infrastructure options for that evaluation. Use the same engine, data, recall target and operating requirements on each candidate. Nirvana's ClickHouse and Elasticsearch results do not substitute for vector-search evidence.
Common questions
How much IOPS does a vector database need? Measure the actual storage operations at the required search quality and query rate. Include ingestion and maintenance. There is no defensible universal number.
Is block storage a replacement for every storage component? No. The engine's architecture determines the required storage services. In particular, a distributed design may also require object storage and other components.
Should I start with pgvector or a dedicated vector database? If PostgreSQL already owns the application data, pgvector is worth evaluating. A dedicated engine may fit different scale, retrieval or operational requirements. Compare with the same workload and correctness criteria.
How do I test Nirvana? Read the LangChain benchmark and Nirvana docs. Prepare a representative dataset, recall target, query mix and recovery objective. Test per-query latency and overall workload completion separately, including both low-concurrency and production-concurrency runs.
Related Posts

How to Choose Kubernetes Storage for High IOPS Workloads
A practical comparison of cloud volumes, NKS with ABS, OpenEBS, Rook/Ceph and Portworx, with a test plan for sustained performance and application recovery.

Which Cloud Is Best for Long-Running AI Agents?
Choose agent hosting by the work it preserves. Compare persistence, recovery, operating models and cost, with clearly scoped Nirvana benchmark evidence.

What Is a Harness? The Agent, The Model, the Harness, and the Sandbox
An agent harness turns a model's decisions into action: it runs the loop, executes each tool call, keeps the notes and calls time. The word is new, the thing is not. What changed in 2026 is that OpenAI, Vercel and LangChain started selling it.