Nirvana Labs×Quirq
Agent sandbox benchmark

Kill the sandbox mid-write. Nirvana came back with 100% of the work instantly.

Nirvana recovered all 5,501 committed checkpoints on a brand-new pod after the compute was killed mid-write, sustains 18,733 durable checkpoints/s at matched 8 vCPU, releases idle compute in 1.5 s, and held 2.71 GiB/s flat through a 30-minute run.

PlatformRecoveredReleased in
Nirvana5,5011.46 s
GKENot counted9.57 s
E2B0

5,501 checkpoints · 20 GiB workspace

Platforms
Nirvana · GKE · E2B
Orchestration
Quirq control plane
Workspace
20 GiB persistent volume
Compute
8 vCPU (Nirvana, GKE)

Background

Our POC partner Quirq put Nirvana, GKE and E2B through the same tests to find out: can the machine under an AI agent be replaced without losing committed work? When it can’t, you pay twice — to redo the lost work, and for idle machines nobody dares release. Put simply: can fast disk unlock cheap agents?

01 · Crash recovery
100%recovered

5,501 of 5,501 checkpoints survived an ungraceful kill. E2B: 0 of 5,000.

02 · Checkpoint rate
GKE

18,733 vs 3,148 commits/s at a matched 8 vCPU.

03 · Pause latency
~1.5 sp50

vs 9.57 s on GKE. Idle over ~10 s is worth releasing.

Net result · cost
cheaper per checkpoint

~$0.004 vs ~$0.024 per 1M durable checkpoints, at an equal vCPU-hour rate.

Requirements

What idempotent compute requires

Kill the machine, hand the agent a fresh one, and it picks up exactly where it left off

01

Committed work lives outside the machine

Repos, checkpoints, caches, databases — on networked storage, never local disk.

02

Durable writes are fast enough to use continuously

Single-digit millisecond checkpoints. Slower gets skipped, and skipped means lost.

03

Releasing and recreating compute is cheap

Pause in ~1.5 s and idle compute is disposable. At 30 s, you pay to leave it running.

fio

Agent sandbox storage characterisation

fio with libaio/O_DIRECT · 20 GiB workspace volume · direct fsync-barrier probe on ABS, both GKE disk tiers, and E2B local disk

Durable fsync barrier?
Nirvana
Yes
GKE pd-balanced
Yes
GKE pd-ssd
Yes
E2B local
No — a no-op*
4K random-write IOPS
Nirvana
~97,700
GKE pd-balanced
6,290
GKE pd-ssd
6,771
E2B local
42,000
Sequential throughput
Nirvana
2.7 GiB/s
GKE pd-balanced
0.28 GiB/s
GKE pd-ssd
0.24 GiB/s
E2B local
0.42 GiB/s
fsync p50 (durability barrier)
Nirvana
1.65 ms
GKE pd-balanced
2.9 ms
GKE pd-ssd
3.0 ms
E2B local
~0 ms (skipped)
fsync p99 / p99.9
Nirvana
2.3 ms / 3.4 ms
GKE pd-balanced
not measured
GKE pd-ssd
not measured
E2B local
not measured
Durable checkpoints/s (fio-derived)
Nirvana
~18,733
GKE pd-balanced
2,157
GKE pd-ssd
2,209
E2B local
not durable
Sustained over 30 min
Nirvana
2.71 GiB/s, flat
GKE pd-balanced
not measured
GKE pd-ssd
not measured
E2B local
not measured
Survives host failure
Nirvana
Yes
GKE pd-balanced
Yes
GKE pd-ssd
Yes
E2B local
No

* E2B in our tested configuration: default template, local disk. Tail percentiles and the 30-minute run were measured on Nirvana only.

Why pd-ssd barely moved

GKE disk IOPS scale with volume size and vCPU, not disk class — at 20 GiB both tiers hit the same ceiling, and unlocking pd-ssd means over-provisioning hundreds of gigabytes a workspace never uses. ABS delivers ~98K IOPS on the same 20 GiB.

No throttling cliff

ABS held 2.71 GiB/s through the final quartile of a 30-minute run, within 0.6% of the first. No burst credits to expire.

E2B’s fsync is a no-op

Tested three ways — os.fsync, an O_DSYNC file, fio with fsync=1 — it returned in ~0 ms every time, no slower than an unsynced write. Fast commits, no durability barrier.

Crash test 1

Crash test

Durable checkpoints written, compute killed ungracefully mid-checkpoint, environment recreated on a brand-new pod (UID verified changed)

Nirvana Agent Sandbox5,501 / 5,501

checkpoints recovered

A new pod attached the same ABS volume and continued exactly where the old one stopped.

GKENot counted

durable, but not run at this count

Persistent disk is a real fsync barrier and the volume reattaches, so GKE survives this failure mode. Its crash durability was verified in the build test (Crash test 2), not with a checkpoint count.

E2B0 / 5,000

checkpoints recovered

In the tested default local-disk configuration, terminating the sandbox removed the workspace with it. We did not test E2B’s pause/resume or persistent snapshots as a recovery mechanism.

Crash test 2

Mid-build chaos

200-module C build plus one final link target, 201 objects total (~4 min) with deterministic checksum · compute destroyed repeatedly mid-build · incremental object cache on the durable volume

Modules already cached on the ABS-backed volume after each destructionNirvana run — each interruption created a brand-new pod
Start
0modules
After crash 1
38modules
After crash 2
76modules
After crash 3
111modules
After crash 4
148modules
After crash 5
186modules
Final
201modules
make resumed from exactly where it stopped, every time · final binary byte-correct

The curve is the Nirvana run — the only one where a per-crash cache count is meaningful, and the schedules differed (five ungraceful destructions on Nirvana, two on E2B). The comparable measure is what an interruption costs: ~10 s on Nirvana to resume into a new pod, ~35 s of pod rebuild and bootstrap on GKE. E2B’s two failures destroyed 52 and 60 modules, each restarting from zero.

Graceful pause — time cost
Nirvana
~1.5 s
GKE
9.57 s
E2B
~0 s (memory snapshot)
Crash — cost per interruption
Nirvana
~10 s to resume into a new pod
GKE
~35 s of rebuild and bootstrap
E2B
full restart from zero
Graceful pause — work lost
Nirvana
0
GKE
0
E2B
0
Ungraceful crash — work lost
Nirvana
0 (cache on ABS)
GKE
0 (cache on PD)
E2B
All — disk dies with VM
Final binary correct?
Nirvana
Yes, across 5 rebuilds
GKE
Yes
E2B
Pause yes / crash no
5 destructions, 0 wasted compiles

GKE survives crashes too, but pays ~35 s of rebuild and bootstrap per crash against ~10 s to resume into a new pod · E2B loses the cache on termination

Scale test

Durable checkpoints at scale

1K → 100K checkpoints · synchronous=FULL, 2 KB + 2 KB/step payload · every commit durably flushed before it counted · matched 8 vCPU · application harness, so rates differ from the fio-derived figures above

Durable checkpoints/s (application harness)
Nirvana· winner
18,733
GKE
3,148
End-to-end duration
Nirvana· winner
5.3 s
GKE
31.8 s
Commit p50
Nirvana· winner
1.1 ms
GKE
3.6 ms
Commit p99
Nirvana· winner
12.4 ms
GKE
72.0 ms
Relative durable-commit throughputThe gap widens with load
1K checkpoints~3× more commits/s
ABS
GKE
E2B
0 durable commits
100K checkpoints18,733 vs 3,148 commits/s — ~6×
ABS
GKE
E2B
0 durable commits
E2B posts the fastest raw commits of the three — 0.04 ms — but its fsync is a no-op, so none of them are durable. Its bar is zero by definition, not by omission.

Lifecycle latency

Absolute percentiles across ~96 runs of create / pause / resume / teardown

Create → running
Nirvana p50
20.6 s
p95
39.6 s
p99
42.5 s
GKE (ref)
~12 s
Pause — release compute
Nirvana p50
1.46 s
p95
1.50 s
p99
1.69 s
GKE (ref)
9.57 s
Teardown
Nirvana p50
0.73 s
p95
0.75 s
p99
0.80 s
GKE (ref)
~30 s
Resume → first useful op
Nirvana p50
8.9 s
p95
9.9 s
p99
10.9 s
GKE (ref)
~12 s+

E2B is not in this table because its lifecycle is a different shape, not because it failed: pause snapshots filesystem and memory, stops compute billing, and resumes in ~0 s. What it does not do is detach a durable volume that a replacement machine can reattach — so it has no create-to-running or teardown phase to compare. Its pause behaviour is in section 04.

Tight tails

Pause holds at ~1.5 s from p50 through p99 (1.46 s → 1.69 s) and teardown under 0.8 s. Provisioning a brand-new workspace costs more (20.6 s p50) — expected for a durable networked volume — but agents pause and resume far more often than they provision.

Warm resume

Across 25 cycles, once the volume reattaches the agent is back at work in milliseconds: git status over 2,394 files completes in ~20 ms cold. The agent doesn’t rebuild its workspace — it continues.

Break-even idle window

At 1.46 s pause p50, any idle gap longer than ~10 s makes pausing profitable. At GKE’s 9.57 s, the window is ~21 s — so a large share of real idle time is only reclaimable on Nirvana.

2,800vs 8,000 compute-hours

A 1,000-agent fleet at ~70% reclaimable idle pays for roughly 2,800 compute-hours on Nirvana instead of 8,000.

Nirvana
~10sidle window
GKE
~21sidle window
Economics

Cost per unit of durable work

Compute-time only, at an assumed equal vCPU-hour rate · excludes storage, cluster management and platform fees · no vendor discounts applied · E2B has no fair denominator, its commits are not crash-durable

Nirvana
Durable checkpoints/s @100K
18,733
Cost per 1M durable checkpoints
~$0.0040
GKE (e2-standard-8, public rate)
Durable checkpoints/s @100K
3,148
Cost per 1M durable checkpoints
~$0.024
E2B
Durable checkpoints/s @100K
Non-durable
Cost per 1M durable checkpoints
Undefined
~6× cheaper per durable unit of work

At equal compute pricing, before any Nirvana discount

What this buys you

The cheaper number is the smaller half of the story. Once compute is genuinely disposable, a persistent agent no longer needs a persistent VM — so the same cluster carries more agents, reclaims idle faster, and completes more jobs per dollar.

Higher agent density

State lives on ABS, not in a VM held open for hours — one node hosts many agents in sequence.

Aggressive idle reclamation

A 10-second idle gap is worth reclaiming at a ~1.5 s pause; GKE needs ~21 s. Most agent idle sits between the two.

Lower cost per completed job

More concurrent agents, faster reclamation, no recomputed work — more finished jobs per cluster-hour.

Comparison

Three agent sandboxes, three definitions of “persistent”

E2B
What “persistent” means
Sandbox state survives a clean pause (memory snapshot)
Survives a crash?
No — local disk dies with VM
Idempotent compute?
No
GKE
What “persistent” means
Workspace state on durable networked PD survives node failure
Survives a crash?
Yes
Idempotent compute?
Yes — 9.57 s pause
Nirvana
What “persistent” means
Workspace state on durable networked ABS survives node failure
Survives a crash?
Yes
Idempotent compute?
Yes — ~1.5 s pause
E2B

The fastest way to run a sandbox. It is not idempotent compute — a crash loses everything.

GKE

Idempotent compute in principle. In practice, ~35 s interruptions and ~30 s teardowns make it expensive to actually treat compute as disposable.

Nirvana Agent Sandboxes

Idempotent compute in practice: crash-durable storage, ~1.5 s pause and sub-second teardown, and a disk fast enough that continuous checkpointing is the default, not an optimization.

Summary

Scorecard

Durable writes (real fsync barrier)
1.65 ms p50 — real barrier, not a no-op
Fast enough for continuous checkpointing
18,733 durable commits/s at 100K — ~6× GKE
Work survives an ungraceful crash
5,501 / 5,501 checkpoints recovered on a new pod
Releasing compute is cheap
1.46 s pause p50 / 0.73 s teardown p50
Real work survives repeated compute loss
200-module build across 5 crashes — 0 wasted compiles
Cost per durable unit of work
~$0.004/M checkpoints vs ~$0.024 on GKE
Methodology

Methodology

Quirq provides the environment and orchestration layer used to run the same agent workloads across each infrastructure configuration. Same workspace image, repository, and dependency installation throughout.

Lifecycle timing
~96 runs across create / pause / resume / teardown; p50, p95, p99 reported
Storage benchmark
fio, 4K random write, libaio + O_DIRECT, iodepth 64, 60 s time-based, group reporting; run at numjobs 1 and 4 and again for 300 s. One pod, one node, one 20 GiB ext4 volume mounted at /home/coder (ABS storageClass on Nirvana, workspace-balanced = pd-balanced on GKE). Direct fsync-barrier probe on ABS, both GKE disk tiers, and E2B local disk
Sustained throughput
30-minute continuous write on the ABS volume; final quartile compared against the first
Durable checkpoints
1K → 100K, synchronous=FULL, 2 KB + 2 KB/step payload; matched 8 vCPU across Nirvana and GKE
Crash recovery
Run on Nirvana and E2B only. Nirvana: kill -9 the process mid-checkpoint, then rebuild the pod under a new UID so the volume reattaches; count of committed state recovered. E2B: sbx.kill() terminates the microVM and its local overlay disk — no durable volume to reattach (E2B Volumes is private beta and was not tested). GKE’s crash durability is demonstrated in the build test instead
Mid-build chaos
200-module C build plus final link target (201 objects), 5 compute destructions on Nirvana and 2 on E2B; byte-correct output verified
Workspace continuity
2,394-file repo + npm ci; directory tree hash, Git HEAD, and dependency fingerprints validated post-resume
Warm resume cycles
25 cycles; time to git status (first useful operation) recorded
Cost model
Compute-time per durable checkpoint at an assumed equal vCPU-hour rate, derived from measured throughput. Compute only — storage, cluster management and platform fees excluded. No vendor discounts applied

Workspace: 20 GiB persistent volume (ABS storageClass:"" on Nirvana; networked persistent disk on GKE). E2B used the same workspace image built as a template. Nirvana: nirvanalabs.io · Quirq: quirq.ai

Quirq
What is Quirq?

Quirq provides the environment and orchestration layer used to run the same agent workloads across each infrastructure configuration.

quirq.ai

Persistence is what lets you run more, faster agents at once.

Persistent agent workspaces on high-IOPS block storage, orchestrated end to end with Quirq.

20 GiB persistent volume · orchestrated with Quirq · August 2026

Powering AI, blockchain, and
databases

Talk to Sales