Back to Blog
EducationAI WorkloadsCloud Infrastructure

How to Choose Cloud Infrastructure for a Vector Database

April WongApril Wong
6 min read
How to Choose Cloud Infrastructure for a Vector Database

Choose cloud infrastructure for a vector database by matching the engine's memory model, search-quality target, query concurrency and recovery requirements. There is no universal IOPS requirement for vector search. A mostly memory-resident index may be CPU- or memory-bound; disk-backed vectors, filtered retrieval, ingestion and index rebuilding can place substantial demand on storage.

The first useful question is where time goes in your own retrieval pipeline. A faster disk cannot remove latency spent creating an embedding, reranking results or waiting for an external model.

A limitation to know before choosing Nirvana: Our LangChain benchmark found slower per-query Qdrant tail latency on Nirvana ABS than on AWS io2, even though Nirvana completed the combined agent workload sooner. A high IOPS figure or faster overall agent run is not evidence of the fastest vector search. If search latency is your main requirement, evaluate that result separately.

Separate database choice from hosting choice

Qdrant, Weaviate, Milvus and pgvector make different architectural choices. You can compare their search capabilities and operating model before deciding where to run them.

EngineInfrastructure considerationsA useful starting point
QdrantMemory placement for vectors, indexes and payloads changes the RAM and disk tradeoffMeasure the selected memory tier and quantization settings on your dataset
WeaviateIts HNSW index requires memory; compression and the selected index type affect resource useSize the index and concurrent workload before assuming more disk performance is the answer
MilvusA distributed architecture separates compute and storage responsibilitiesInclude the deployed version's full component and storage requirements in the plan
PostgreSQL with pgvectorVector search shares a database with relational data, writes, maintenance and SQL queriesCheck index behavior and contention alongside normal PostgreSQL operations

Primary references: Qdrant memory tiers, Weaviate resource planning, Milvus architecture and pgvector documentation.

A managed database transfers much of the operational work to the service provider. Self-hosting gives you more control over placement, configuration and infrastructure, while leaving upgrades, replication, monitoring and recovery with your team. A cloud VM is not automatically a managed vector database.

Size the data before sizing the machine

As an illustrative lower bound, ten million vectors with 768 dimensions stored as float32 values contain 30.72 GB of raw vector values: 10,000,000 × 768 × 4 bytes. That is about 28.6 GiB. It excludes index structures, payloads, metadata, replication, logs, temporary build space and operating-system overhead.

Do not turn that arithmetic into a RAM recommendation. The index type, engine implementation, compression and cache behavior determine what must remain in memory. Leave room for ingestion and rebuilds as well as steady-state queries.

Qdrant's storage documentation describes configurable memory placement. Weaviate's guidance describes its HNSW memory requirements. Those are reasons to test engine-specific settings, rather than using one capacity multiplier for every database.

Find the bottleneck in the retrieval pipeline

Break an end-to-end request into embedding generation, network transit, database search, filtering or payload reads, reranking and answer generation. Measure each stage separately and under realistic concurrency.

If database search slows while CPU is saturated, examine query and index settings before buying more IOPS. If memory pressure causes frequent disk access and storage queues grow, a faster storage path or a different memory layout may help. If search is fast but the full request is slow, focus on the stage actually consuming time.

Keep the required search quality fixed. Recall measures how many relevant nearest neighbours the approximate search finds compared with the chosen reference. A faster result at substantially lower recall is not an equivalent result. Include real filters and tenant distributions; an unfiltered demonstration can miss the expensive production path.

Use IOPS as a measured requirement

Observe storage reads and writes while the database serves the target query rate and ingestion load. Record block sizes, cache misses, queue depth, latency and CPU utilization. Repeat after a restart and with a working set larger than effective memory.

You can use a simple planning identity: storage operations per second are approximately queries per second multiplied by the average storage operations required per query, plus writes and background work. The operations-per-query term must come from measurement. It changes with caching, filters, index layout and data distribution.

This explains why neither vector count nor a provider's peak IOPS number can determine application performance on its own. The meaningful target is the required query rate and recall at the agreed p95/p99 latency.

Queue depth is the number of outstanding storage operations at the layer being measured. At QD=1, that I/O stream waits for one operation to finish before issuing another. Its performance depends heavily on individual-operation latency. A high-queue-depth throughput test measures a different situation, with many requests in flight.

The LangChain report describes its per-query Qdrant tradeoff as QD=1. Its raw 4 KB random-read fio test uses QD=256. Those numbers cannot be substituted for each other. A search that follows dependent reads through an HNSW graph may have limited I/O parallelism along that path.

Batching and concurrent queries can create more parallel work, but do not automatically shorten a single chain of dependent reads. Many agents can also create aggregate device concurrency while individual query paths remain latency-sensitive. Measure the queue depth actually reached; do not assume every vector engine or query always runs at QD=1.

Compare infrastructure with one test plan

Use the same embedding set, distance metric, search-quality target, filters and query distribution on each option. Pin the engine version and index settings. Record both warm behavior and a defined cold or recovery state.

Test single-query latency and throughput at realistic concurrency separately. Record actual storage queue depth. Measure query throughput, p50/p95/p99 latency, recall, ingestion rate, time until new data is searchable, index-build time and storage consumed. Run queries while ingestion and maintenance continue. Report failures and timeouts alongside completed requests.

Then replace a worker or restore from backup in a test environment. Confirm the dataset and index are correct and record how long it takes to return to the service objective. Include replica placement and the failure boundary being tested; a successful local restart does not prove multi-region disaster recovery.

Price the smallest configuration that passes. Include replicas, memory, compute, block and object storage where required, snapshots, backups, network transfer and operational effort. Useful comparisons are cost at a fixed dataset and query rate, or cost per successful query at a fixed recall and latency target. A bare monthly VM price leaves too much out.

What Nirvana's LangChain benchmark actually shows

The published LangChain read benchmark compares a mixed Qdrant, Redis and PostgreSQL workload across Nirvana ABS and four AWS storage configurations. The R3 production-scale run uses five million 768-dimensional vectors across 50 collections, Qdrant 1.16 with inline storage and INT8 scalar quantization, and 1,000 agents completing 100 tasks each. Each task includes two operations per service.

Metric at 100,000 tasks, R3Nirvana ABSAWS io2-64k
Qdrant vector-search p99301 ms140 ms
Complete agent-task p99725 ms771 ms
Total workload completion58 minutes69 minutes

Lower is better in every row. Nirvana's Qdrant p99 was about 2.15 times the io2 result in this run. The complete workload nevertheless finished about 16% sooner. These are different outcomes: a throughput advantage across the mixed workload does not establish a per-query vector-search advantage. Per-service p99 values also cannot simply be added to reconstruct task p99.

The configurations had the same nominal vCPU count, RAM capacity and volume size, but used different compute platforms. This is a comparison of the tested deployments, not an experiment isolating storage batching as the sole cause. It does not establish the same result for Weaviate, Milvus, pgvector or another Qdrant configuration.

When to evaluate Nirvana

For an application whose main requirement is low tail latency on individual vector searches, the published result favors the tested io2 configuration. Start with query latency at your required recall and concurrency; do not choose ABS on peak IOPS alone.

For mixed agent workloads where overall completion time and sustained throughput matter, the LangChain result provides a reason to evaluate Nirvana. Confirm that the vector-search component still meets your application's latency objective. Ingestion, index builds and recovery need their own measurements; this result does not establish a Nirvana advantage for them.

ABS and NKS are infrastructure options for that evaluation. Use the same engine, data, recall target and operating requirements on each candidate. Nirvana's ClickHouse and Elasticsearch results do not substitute for vector-search evidence.

Common questions

How much IOPS does a vector database need? Measure the actual storage operations at the required search quality and query rate. Include ingestion and maintenance. There is no defensible universal number.

Is block storage a replacement for every storage component? No. The engine's architecture determines the required storage services. In particular, a distributed design may also require object storage and other components.

Should I start with pgvector or a dedicated vector database? If PostgreSQL already owns the application data, pgvector is worth evaluating. A dedicated engine may fit different scale, retrieval or operational requirements. Compare with the same workload and correctness criteria.

How do I test Nirvana? Read the LangChain benchmark and Nirvana docs. Prepare a representative dataset, recall target, query mix and recovery objective. Test per-query latency and overall workload completion separately, including both low-concurrency and production-concurrency runs.

Powering AI, blockchain, and
databases

Talk to Sales