Back to Blog
EducationDatabase

How to Test Redis AOF Rewrites and Recovery Before Production

April WongApril Wong
5 min read
How to Test Redis AOF Rewrites and Recovery Before Production

A Redis instance can serve normal traffic well and still miss its latency or recovery targets during an AOF rewrite. Before production, test the maintenance and failure conditions your application must survive. This guide builds a repeatable acceptance test for Redis AOF (Append Only File) persistence, using synthetic data in an isolated environment.

TL;DR

  • Agree on latency, acknowledged-write loss and recovery-time limits before testing; keep the persistence policy fixed across infrastructure candidates.
  • Measure the same workload before, during and after a completed AOF rewrite, including memory pressure and delayed fsyncs.
  • Keep an independent record of acknowledged writes. A successful PING or a matching key count does not prove correct recovery.
  • Test backup restoration separately from restarting on the existing volume, and use the backup procedure for your Redis version.

What must your Redis deployment survive?

Start with three application requirements: p99 latency at the intended request rate, the maximum acceptable loss of acknowledged writes, and the time allowed to restore useful service. A rebuildable cache and Redis holding workflow state can have very different acceptance criteria.

Record the Redis version, dataset size, value sizes, expiry behavior, eviction policy and client location. Repeat tests at expected peak dataset size, not just on a small freshly populated instance. Keep the load generator and its logs outside the Redis host so a host failure does not erase your evidence.

Use these as separate test cases, with a pass/fail decision for each:

ConditionEvidence to retain
Steady trafficOffered and completed rate, latency percentiles and errors
AOF rewrite under trafficRewrite interval, application tail latency, memory and disk headroom
Restart or abrupt failureFailure type, missing acknowledged writes and time to useful service
Independent restoreBackup boundary, restored values and application checks

Which persistence settings should you verify?

For a fresh test deployment, a common starting configuration is:

appendonly yes
appendfsync everysec

This is only the persistence fragment. Configure the intended durable data directory, private connectivity, authentication and memory limits separately. If an existing instance currently uses RDB snapshots, follow Redis’s documented conversion procedure; changing a file and restarting is not a safe migration shortcut.

With everysec, a successful write reply is not confirmation that that write has been fsynced. The always policy synchronizes appended batches before replies; no leaves flushing to the operating system. Evaluate each policy against its own durability requirements.

From an authenticated administrative CLI session, capture the effective configuration:

CONFIG GET appendonly
CONFIG GET appendfsync
CONFIG GET no-appendfsync-on-rewrite
CONFIG GET dir
CONFIG GET appenddirname
CONFIG GET maxmemory
CONFIG GET maxmemory-policy
INFO persistence

For a baseline that keeps AOF synchronization active during rewrites, verify no-appendfsync-on-rewrite is no. Redis’s latency guide explains the durability trade-off of changing it. Save settings with every result so a faster run cannot silently reflect weaker persistence.

How do you measure a rewrite under realistic load?

Reproduce the application’s command mix, payload distribution, concurrency and pipeline depth. Redis’s benchmarking guide explains why client and network choices can dominate results. A SET-only throughput test is a useful component check, but it does not represent a multi-command application transaction.

Warm the dataset, run a stable baseline, then request one rewrite while traffic continues:

BGREWRITEAOF
INFO persistence

The command response can mean the rewrite was scheduled behind an existing background save. Record its actual start and completion; do not time only the API response.

Sample INFO persistence throughout the run. The INFO reference defines the relevant fields:

  • aof_rewrite_in_progress and aof_rewrite_scheduled: distinguish active work from queued work.
  • aof_last_bgrewrite_status and aof_last_write_status: check successful completion and ongoing writes.
  • aof_last_cow_size: inspect copy-on-write memory associated with the completed rewrite.
  • aof_delayed_fsync: compare counter changes across the test, accounting for restarts.

Correlate those timestamps with application p50/p95/p99, errors, host memory, CPU and device latency. A fork pause, memory pressure and storage contention require different remedies. Retain the recovery period after the rewrite too. Repeat the full sequence under the same conditions rather than selecting the best run.

How do you prove that acknowledged writes survive?

Use a small correctness workload alongside the performance workload. Assign each write a unique key and deterministic value. After receiving its successful reply, record the key, expected value and acknowledgement time in an independent ledger. Do not overwrite these validation keys or expire them during this test; exercise TTL behavior in a separate case.

Run a clean restart first. Then test an approved abrupt failure in the isolated environment. Name the fault precisely: terminating the Redis process leaves the operating system alive, so it does not establish what would survive loss of the host or its volatile caches.

After recovery, compare every acknowledged validation record with its expected value. Report missing records, mismatches and the interval between their acknowledgement and the fault. Keep requests with unknown outcomes separate: a lost reply does not establish whether Redis applied the write.

Measure recovery from the failure timestamp until the application reconnects and completes its required read/write checks. Record server loading, client retries and backlog recovery separately. A responsive port or PING is only an intermediate milestone.

What makes the deployment ready for production?

Restore a backup into a separate instance and repeat the value checks. Restarting against the original volume proves a different recovery path. For Redis 7 and later, AOF includes a manifest and multiple files; copying one apparent log file is insufficient. Follow the version-appropriate backup procedure, including its rules around concurrent rewrites.

Approve the deployment only when each test meets its agreed latency, loss and recovery limits. Keep the configuration, dataset generator, results and restore steps together. Include headroom and alerts for persistence errors, disk capacity and memory pressure, plus a named operator for recovery.

For interpreting storage symptoms, read When Does Redis Need Faster Storage? This guide supplies a test method; it does not report a new Redis benchmark or extrapolate a combined agent-workload result into standalone Redis performance.

Plan a Redis infrastructure evaluation with Nirvana Labs. Bring your dataset size, persistence policy, peak write rate and recovery targets.

FAQ

Does a persistent volume provide Redis high availability?

It preserves storage independently of a process. Availability also depends on the deployment architecture, failure domains and client behavior. Test failover separately if replicas or Sentinel are part of your design.

Can I compare everysec with always?

Yes, as different durability configurations. Label the policies explicitly and choose the one that meets the application’s acknowledged-write requirements.

Why can a key-count check miss recovery errors?

Two datasets can contain the same number of keys while holding different values or missing different records. Validate the acknowledged keys and their contents.

About Nirvana Labs

Nirvana Labs is a high-performance storage cloud purpose built for blockchain, AI and databases i.e. the most demanding, real-time, stateful workloads. Accelerated Block Storage (ABS) offers 20K baseline IOPS included, no over provisioning. Nirvana Kubernetes Service (NKS) with Karpenter auto-scaling, high clock-speed compute and private networking. Backed by Jump Trading, Crucible, etc with 50+ customers live in production today.

Learn more at Nirvana Labs

Nirvana Cloud | Pricing | Blog | Docs | Changelog | LinkedIn | Twitter | Telegram | YouTube

Powering AI, blockchain, and
databases

Talk to Sales