Back to Blog
EducationDatabase

How to Benchmark PostgreSQL WAL and Commit Performance on Cloud Storage

April WongApril Wong
5 min read
How to Benchmark PostgreSQL WAL and Commit Performance on Cloud Storage

A PostgreSQL storage evaluation should show how transactions behave at the application’s required load, with its durability settings intact. This guide builds a repeatable test for PostgreSQL WAL (Write-Ahead Log), commit waits and checkpoint activity on cloud block storage. The examples target PostgreSQL 18 and are test procedures, not newly measured Nirvana benchmark results.

TL;DR

  • Match durability, compute, dataset and client placement before comparing storage; synchronous replicas can also determine commit latency.
  • Run a controlled pgbench workload, then an application-representative transaction test through checkpoint activity.
  • Keep transaction latency, COMMIT response time and WAL synchronization time separate. They measure different parts of the path.
  • Use PostgreSQL 18’s version-specific statistics, retain raw latency samples, and prove backup recovery before migration.

What should stay constant in a storage comparison?

Create an isolated test environment and record the PostgreSQL and client versions, CPU model, vCPU count, RAM, filesystem, volume size, IOPS/throughput limits and client location. Identify where both the data directory and WAL reside. A separate WAL volume changes the configuration being evaluated.

Confirm that the intended persistent volume is mounted after reboot and that PostgreSQL starts against it. Keep dataset size, schema, indexes and maintenance settings consistent across candidates. Record any differences you cannot eliminate.

From the same database role and connection path used by the workload, capture:

SHOW server_version;
SHOW fsync;
SHOW full_page_writes;
SHOW synchronous_commit;
SHOW synchronous_standby_names;
SHOW wal_sync_method;
SHOW checkpoint_timeout;
SHOW max_wal_size;

Keep fsync and full_page_writes enabled for the durable baseline. With synchronous standbys configured, synchronous_commit=on also waits for the required standby flushes. Turning it off changes acknowledged-transaction durability; disabling fsync can risk corruption. PostgreSQL’s WAL configuration reference documents these distinctions.

How do you build a repeatable pgbench experiment?

Point PostgreSQL client connection settings at the isolated server before running anything. The following shell sequence creates a dedicated database and stops if creation fails. Initialization replaces standard pgbench tables, so never aim it at an application database.

set -eu
createdb nirvana_pgbench_lab
pgbench -i -s 50 nirvana_pgbench_lab
pgbench -c 8 -j 4 -T 60 nirvana_pgbench_lab
pgbench -c 8 -j 4 -T 900 -P 10 -l   --log-prefix=run01 nirvana_pgbench_lab

The first workload run is a warm-up; the second records individual transactions. With -j 4 and -l, pgbench writes one log file per worker thread: run01.<pid>, run01.<pid>.1, run01.<pid>.2 and run01.<pid>.3. Combine the individual transaction samples from all four files for that run before calculating percentiles. Scale 50 and eight clients are illustrative starting points, not a storage-sizing recommendation. Inspect table size relative to memory. A cache-resident experiment and a larger working set answer different questions.

Use pgbench to compare identical experiments, then replace its built-in transaction mix with the application’s actual writes, indexes, batch sizes and connection behavior. Test increasing concurrency and an explicit target arrival rate with -R. Under rate limiting, reported latency includes schedule lag; retain that lag as evidence of an overloaded system or load generator.

Repeat each condition, restore equivalent starting data, and retain every run. Check that the measurement window actually includes checkpoint activity; elapsed duration alone cannot prove it. Record warm-up policy and run order so caching or a progressively changed dataset does not masquerade as a storage improvement.

Which latency are you actually measuring?

Write the measurement boundary beside every percentile. Three commonly confused values are:

MetricBoundaryUseful for
Transaction latencyClient-observed completion of the whole transaction; include arrival scheduling when applicableApplication responsiveness at the tested load
COMMIT response timeClient sends COMMIT to client receives its replyCommit-path investigation, including network and replication waits
WAL synchronization timeDatabase-recorded time in the relevant synchronization operationOne component of durability-related I/O

Calculate p50, p95 and p99 from individual successful transaction samples; publish failures, retries and skipped transactions alongside them. Keep raw per-transaction logs when percentiles matter. Aggregate summaries and average statement timings cannot reconstruct a commit p99.

To investigate COMMIT specifically, instrument its client send/reply interval in a representative transaction harness. Use the same clock for each interval and record connection setup separately. This remains a client-observed measurement, with more in it than device service time.

pg_test_fsync can help examine synchronization methods on the intended filesystem in an isolated test location. Its result is a component diagnostic; it does not establish application throughput or tail latency.

How do you connect slow transactions to WAL and checkpoints?

Take before-and-after statistics snapshots in separate transactions and compare deltas only when their reset timestamps match. For PostgreSQL 18:

SELECT wal_records, wal_fpi, wal_bytes,
       wal_buffers_full, stats_reset
FROM pg_stat_wal;

SELECT backend_type, object, context,
       writes, write_time, fsyncs, fsync_time, stats_reset
FROM pg_stat_io
WHERE object = 'wal';

SELECT num_timed, num_requested, num_done,
       write_time, sync_time, buffers_written, stats_reset
FROM pg_stat_checkpointer;

These are PostgreSQL 18 queries. Older releases expose different fields. In 18, WAL timing is in pg_stat_io; pg_stat_wal supplies WAL generation counters. Check track_wal_io_timing before interpreting timing values, and record measurement overhead. The statistics reference explains the views, timing controls and caching limitations. In PostgreSQL 18, num_timed and num_requested include skipped checkpoints; use num_done to count only completed checkpoints.

Sample wait events from pg_stat_activity, and align them with client latency, CPU, memory, device queues and replica health. A lock wait, saturated client and synchronous-replica wait can each increase latency without proving a primary-storage bottleneck. Cluster-wide counters can include unrelated work, so keep the test cluster quiet.

Compare whole-run percentiles with shorter windows around checkpoint activity. If the overall p99 passes while the checkpoint interval repeatedly fails the application target, report both. Summed I/O timing divided by operation count gives an average for those operations, not a transaction percentile.

What evidence should support a migration decision?

Set acceptance criteria before inspecting the results: the required arrival rate, p99 latency ceiling, error budget, recovery time and acknowledged-data-loss limit. Compare configurations at those requirements, including compute, storage, replicas, backups and operating effort.

Then restore into a separate environment. For point-in-time recovery, establish a usable base backup and a continuous WAL archive covering the required interval, following PostgreSQL’s recovery documentation. Validate known records and useful application requests. pg_verifybackup provides additional checks; it does not replace a restore rehearsal.

Our PostgreSQL commit-latency explainer and the LangChain agent benchmark discuss earlier measurements and their workload limits. Use this procedure to evaluate your own workload on Nirvana compute and Accelerated Block Storage (ABS), preserving the same comparison rules for every candidate.

Plan a PostgreSQL evaluation with Nirvana Labs. Bring the transaction mix, dataset size, durability policy and target load.

FAQ

Does higher pgbench TPS prove faster storage?

It shows more completed transactions for that experiment. CPU, locks, caching, networking and durability settings can change TPS too; inspect the correlated evidence.

Should WAL always use a separate volume?

Treat it as a configuration to test. Separate volumes may change contention, cost and operations; compare the complete deployment under the same workload.

Why can WAL synchronization time differ from COMMIT latency?

The synchronization operation covers one part of the path. A client-observed commit also includes other waits and communication, and concurrent transactions may share durability work.

About Nirvana Labs

Nirvana Labs is a high-performance storage cloud purpose built for blockchain, AI and databases i.e. the most demanding, real-time, stateful workloads. Accelerated Block Storage (ABS) offers 20K baseline IOPS included, no over provisioning. Nirvana Kubernetes Service (NKS) with Karpenter auto-scaling, high clock-speed compute and private networking. Backed by Jump Trading, Crucible, etc with 50+ customers live in production today.

Learn more at Nirvana Labs

Nirvana Cloud | Pricing | Blog | Docs | Changelog | LinkedIn | Twitter | Telegram | YouTube

Powering AI, blockchain, and
databases

Talk to Sales