
Agent Sandbox Benchmark: Nirvana vs GKE vs E2B
Nirvana recovered all 5,501 committed checkpoints on a brand-new pod after the compute was killed mid-write, sustains 18,733 durable checkpoints/s at matched 8 vCPU, releases idle compute in 1.5 s, and held 2.71 GiB/s flat through a 30-minute run.
- Platforms
- Nirvana · GKE · E2B
- Orchestration
- Quirq control plane
- Workspace
- 20 GiB persistent volume
- Compute
- 8 vCPU (Nirvana, GKE)
5,501 of 5,501 checkpoints survived an ungraceful kill. E2B: 0 of 5,000.
18,733 vs 3,148 commits/s at a matched 8 vCPU.
vs 9.57 s on GKE. Idle over ~10 s is worth releasing.
~$0.004 vs ~$0.024 per 1M durable checkpoints, at an equal vCPU-hour rate.
Background
Our POC partner Quirq put Nirvana, GKE and E2B through the same tests to find out: can the machine under an AI agent be replaced without losing committed work? When it can’t, you pay twice — to redo the lost work, and for idle machines nobody dares release. Put simply: can fast disk unlock cheap agents?
What idempotent compute requires
Kill the machine, hand the agent a fresh one, and it picks up exactly where it left off
Committed work lives outside the machine
Repos, checkpoints, caches, databases — on networked storage, never local disk.
Durable writes are fast enough to use continuously
Single-digit millisecond checkpoints. Slower gets skipped, and skipped means lost.
Releasing and recreating compute is cheap
Pause in ~1.5 s and idle compute is disposable. At 30 s, you pay to leave it running.
Agent sandbox storage characterisation
fio with libaio/O_DIRECT · 20 GiB workspace volume · direct fsync-barrier probe on ABS, both GKE disk tiers, and E2B local disk
| Metric | Nirvana | GKE pd-balanced | GKE pd-ssd | E2B local |
|---|---|---|---|---|
| Durable fsync barrier? | Yes | Yes | Yes | No — a no-op* |
| 4K random-write IOPS | ~97,700 | 6,290 | 6,771 | 42,000 |
| Sequential throughput | 2.7 GiB/s | 0.28 GiB/s | 0.24 GiB/s | 0.42 GiB/s |
| fsync p50 (durability barrier) | 1.65 ms | 2.9 ms | 3.0 ms | ~0 ms (skipped) |
| fsync p99 / p99.9 | 2.3 ms / 3.4 ms | not measured | not measured | not measured |
| Durable checkpoints/s (fio-derived) | ~18,733 | 2,157 | 2,209 | not durable |
| Sustained over 30 min | 2.71 GiB/s, flat | not measured | not measured | not measured |
| Survives host failure | Yes | Yes | Yes | No |
- Nirvana
- Yes
- GKE pd-balanced
- Yes
- GKE pd-ssd
- Yes
- E2B local
- No — a no-op*
- Nirvana
- ~97,700
- GKE pd-balanced
- 6,290
- GKE pd-ssd
- 6,771
- E2B local
- 42,000
- Nirvana
- 2.7 GiB/s
- GKE pd-balanced
- 0.28 GiB/s
- GKE pd-ssd
- 0.24 GiB/s
- E2B local
- 0.42 GiB/s
- Nirvana
- 1.65 ms
- GKE pd-balanced
- 2.9 ms
- GKE pd-ssd
- 3.0 ms
- E2B local
- ~0 ms (skipped)
- Nirvana
- 2.3 ms / 3.4 ms
- GKE pd-balanced
- not measured
- GKE pd-ssd
- not measured
- E2B local
- not measured
- Nirvana
- ~18,733
- GKE pd-balanced
- 2,157
- GKE pd-ssd
- 2,209
- E2B local
- not durable
- Nirvana
- 2.71 GiB/s, flat
- GKE pd-balanced
- not measured
- GKE pd-ssd
- not measured
- E2B local
- not measured
- Nirvana
- Yes
- GKE pd-balanced
- Yes
- GKE pd-ssd
- Yes
- E2B local
- No
* E2B in our tested configuration: default template, local disk. Tail percentiles and the 30-minute run were measured on Nirvana only.
GKE disk IOPS scale with volume size and vCPU, not disk class — at 20 GiB both tiers hit the same ceiling, and unlocking pd-ssd means over-provisioning hundreds of gigabytes a workspace never uses. ABS delivers ~98K IOPS on the same 20 GiB.
ABS held 2.71 GiB/s through the final quartile of a 30-minute run, within 0.6% of the first. No burst credits to expire.
Tested three ways — os.fsync, an O_DSYNC file, fio with fsync=1 — it returned in ~0 ms every time, no slower than an unsynced write. Fast commits, no durability barrier.
Crash test
Durable checkpoints written, compute killed ungracefully mid-checkpoint, environment recreated on a brand-new pod (UID verified changed)
checkpoints recovered
A new pod attached the same ABS volume and continued exactly where the old one stopped.
durable, but not run at this count
Persistent disk is a real fsync barrier and the volume reattaches, so GKE survives this failure mode. Its crash durability was verified in the build test (Crash test 2), not with a checkpoint count.
checkpoints recovered
In the tested default local-disk configuration, terminating the sandbox removed the workspace with it. We did not test E2B’s pause/resume or persistent snapshots as a recovery mechanism.
Mid-build chaos
200-module C build plus one final link target, 201 objects total (~4 min) with deterministic checksum · compute destroyed repeatedly mid-build · incremental object cache on the durable volume
make resumed from exactly where it stopped, every time · final binary byte-correctThe curve is the Nirvana run — the only one where a per-crash cache count is meaningful, and the schedules differed (five ungraceful destructions on Nirvana, two on E2B). The comparable measure is what an interruption costs: ~10 s on Nirvana to resume into a new pod, ~35 s of pod rebuild and bootstrap on GKE. E2B’s two failures destroyed 52 and 60 modules, each restarting from zero.
| Interruption type | Nirvana | GKE | E2B |
|---|---|---|---|
| Graceful pause — time cost | ~1.5 s | 9.57 s | ~0 s (memory snapshot) |
| Crash — cost per interruption | ~10 s to resume into a new pod | ~35 s of rebuild and bootstrap | full restart from zero |
| Graceful pause — work lost | 0 | 0 | 0 |
| Ungraceful crash — work lost | 0 (cache on ABS) | 0 (cache on PD) | All — disk dies with VM |
| Final binary correct? | Yes, across 5 rebuilds | Yes | Pause yes / crash no |
- Nirvana
- ~1.5 s
- GKE
- 9.57 s
- E2B
- ~0 s (memory snapshot)
- Nirvana
- ~10 s to resume into a new pod
- GKE
- ~35 s of rebuild and bootstrap
- E2B
- full restart from zero
- Nirvana
- 0
- GKE
- 0
- E2B
- 0
- Nirvana
- 0 (cache on ABS)
- GKE
- 0 (cache on PD)
- E2B
- All — disk dies with VM
- Nirvana
- Yes, across 5 rebuilds
- GKE
- Yes
- E2B
- Pause yes / crash no
GKE survives crashes too, but pays ~35 s of rebuild and bootstrap per crash against ~10 s to resume into a new pod · E2B loses the cache on termination
Durable checkpoints at scale
1K → 100K checkpoints · synchronous=FULL, 2 KB + 2 KB/step payload · every commit durably flushed before it counted · matched 8 vCPU · application harness, so rates differ from the fio-derived figures above
| At 100K durable checkpoints | Nirvanawinner | GKE |
|---|---|---|
| Durable checkpoints/s (application harness) | 18,733 | 3,148 |
| End-to-end duration | 5.3 s | 31.8 s |
| Commit p50 | 1.1 ms | 3.6 ms |
| Commit p99 | 12.4 ms | 72.0 ms |
- Nirvana· winner
- 18,733
- GKE
- 3,148
- Nirvana· winner
- 5.3 s
- GKE
- 31.8 s
- Nirvana· winner
- 1.1 ms
- GKE
- 3.6 ms
- Nirvana· winner
- 12.4 ms
- GKE
- 72.0 ms
Lifecycle latency
Absolute percentiles across ~96 runs of create / pause / resume / teardown
| Phase | Nirvana p50 | p95 | p99 | GKE (ref) |
|---|---|---|---|---|
| Create → running | 20.6 s | 39.6 s | 42.5 s | ~12 s |
| Pause — release compute | 1.46 s | 1.50 s | 1.69 s | 9.57 s |
| Teardown | 0.73 s | 0.75 s | 0.80 s | ~30 s |
| Resume → first useful op | 8.9 s | 9.9 s | 10.9 s | ~12 s+ |
- Nirvana p50
- 20.6 s
- p95
- 39.6 s
- p99
- 42.5 s
- GKE (ref)
- ~12 s
- Nirvana p50
- 1.46 s
- p95
- 1.50 s
- p99
- 1.69 s
- GKE (ref)
- 9.57 s
- Nirvana p50
- 0.73 s
- p95
- 0.75 s
- p99
- 0.80 s
- GKE (ref)
- ~30 s
- Nirvana p50
- 8.9 s
- p95
- 9.9 s
- p99
- 10.9 s
- GKE (ref)
- ~12 s+
E2B is not in this table because its lifecycle is a different shape, not because it failed: pause snapshots filesystem and memory, stops compute billing, and resumes in ~0 s. What it does not do is detach a durable volume that a replacement machine can reattach — so it has no create-to-running or teardown phase to compare. Its pause behaviour is in section 04.
Pause holds at ~1.5 s from p50 through p99 (1.46 s → 1.69 s) and teardown under 0.8 s. Provisioning a brand-new workspace costs more (20.6 s p50) — expected for a durable networked volume — but agents pause and resume far more often than they provision.
Across 25 cycles, once the volume reattaches the agent is back at work in milliseconds: git status over 2,394 files completes in ~20 ms cold. The agent doesn’t rebuild its workspace — it continues.
At 1.46 s pause p50, any idle gap longer than ~10 s makes pausing profitable. At GKE’s 9.57 s, the window is ~21 s — so a large share of real idle time is only reclaimable on Nirvana.
A 1,000-agent fleet at ~70% reclaimable idle pays for roughly 2,800 compute-hours on Nirvana instead of 8,000.
Cost per unit of durable work
Compute-time only, at an assumed equal vCPU-hour rate · excludes storage, cluster management and platform fees · no vendor discounts applied · E2B has no fair denominator, its commits are not crash-durable
| Platform | Durable checkpoints/s @100K | Cost per 1M durable checkpoints |
|---|---|---|
| Nirvana | 18,733 | ~$0.0040 |
| GKE (e2-standard-8, public rate) | 3,148 | ~$0.024 |
| E2B | Non-durable | Undefined |
- Durable checkpoints/s @100K
- 18,733
- Cost per 1M durable checkpoints
- ~$0.0040
- Durable checkpoints/s @100K
- 3,148
- Cost per 1M durable checkpoints
- ~$0.024
- Durable checkpoints/s @100K
- Non-durable
- Cost per 1M durable checkpoints
- Undefined
At equal compute pricing, before any Nirvana discount
What this buys you
The cheaper number is the smaller half of the story. Once compute is genuinely disposable, a persistent agent no longer needs a persistent VM — so the same cluster carries more agents, reclaims idle faster, and completes more jobs per dollar.
State lives on ABS, not in a VM held open for hours — one node hosts many agents in sequence.
A 10-second idle gap is worth reclaiming at a ~1.5 s pause; GKE needs ~21 s. Most agent idle sits between the two.
More concurrent agents, faster reclamation, no recomputed work — more finished jobs per cluster-hour.
Three agent sandboxes, three definitions of “persistent”
| Platform | What “persistent” means | Survives a crash? | Idempotent compute? |
|---|---|---|---|
| E2B | Sandbox state survives a clean pause (memory snapshot) | No — local disk dies with VM | No |
| GKE | Workspace state on durable networked PD survives node failure | Yes | Yes — 9.57 s pause |
| Nirvana | Workspace state on durable networked ABS survives node failure | Yes | Yes — ~1.5 s pause |
- What “persistent” means
- Sandbox state survives a clean pause (memory snapshot)
- Survives a crash?
- No — local disk dies with VM
- Idempotent compute?
- No
- What “persistent” means
- Workspace state on durable networked PD survives node failure
- Survives a crash?
- Yes
- Idempotent compute?
- Yes — 9.57 s pause
- What “persistent” means
- Workspace state on durable networked ABS survives node failure
- Survives a crash?
- Yes
- Idempotent compute?
- Yes — ~1.5 s pause
The fastest way to run a sandbox. It is not idempotent compute — a crash loses everything.
Idempotent compute in principle. In practice, ~35 s interruptions and ~30 s teardowns make it expensive to actually treat compute as disposable.
Idempotent compute in practice: crash-durable storage, ~1.5 s pause and sub-second teardown, and a disk fast enough that continuous checkpointing is the default, not an optimization.
Scorecard
- Durable writes (real fsync barrier)
- 1.65 ms p50 — real barrier, not a no-op
- Fast enough for continuous checkpointing
- 18,733 durable commits/s at 100K — ~6× GKE
- Work survives an ungraceful crash
- 5,501 / 5,501 checkpoints recovered on a new pod
- Releasing compute is cheap
- 1.46 s pause p50 / 0.73 s teardown p50
- Real work survives repeated compute loss
- 200-module build across 5 crashes — 0 wasted compiles
- Cost per durable unit of work
- ~$0.004/M checkpoints vs ~$0.024 on GKE
Methodology
Quirq provides the environment and orchestration layer used to run the same agent workloads across each infrastructure configuration. Same workspace image, repository, and dependency installation throughout.
- Lifecycle timing
- ~96 runs across create / pause / resume / teardown; p50, p95, p99 reported
- Storage benchmark
- fio, 4K random write, libaio + O_DIRECT, iodepth 64, 60 s time-based, group reporting; run at numjobs 1 and 4 and again for 300 s. One pod, one node, one 20 GiB ext4 volume mounted at /home/coder (ABS storageClass on Nirvana, workspace-balanced = pd-balanced on GKE). Direct fsync-barrier probe on ABS, both GKE disk tiers, and E2B local disk
- Sustained throughput
- 30-minute continuous write on the ABS volume; final quartile compared against the first
- Durable checkpoints
- 1K → 100K, synchronous=FULL, 2 KB + 2 KB/step payload; matched 8 vCPU across Nirvana and GKE
- Crash recovery
- Run on Nirvana and E2B only. Nirvana: kill -9 the process mid-checkpoint, then rebuild the pod under a new UID so the volume reattaches; count of committed state recovered. E2B: sbx.kill() terminates the microVM and its local overlay disk — no durable volume to reattach (E2B Volumes is private beta and was not tested). GKE’s crash durability is demonstrated in the build test instead
- Mid-build chaos
- 200-module C build plus final link target (201 objects), 5 compute destructions on Nirvana and 2 on E2B; byte-correct output verified
- Workspace continuity
- 2,394-file repo + npm ci; directory tree hash, Git HEAD, and dependency fingerprints validated post-resume
- Warm resume cycles
- 25 cycles; time to git status (first useful operation) recorded
- Cost model
- Compute-time per durable checkpoint at an assumed equal vCPU-hour rate, derived from measured throughput. Compute only — storage, cluster management and platform fees excluded. No vendor discounts applied
Workspace: 20 GiB persistent volume (ABS storageClass:"" on Nirvana; networked persistent disk on GKE). E2B used the same workspace image built as a template. Nirvana: nirvanalabs.io · Quirq: quirq.ai

Persistence is what lets you run more, faster agents at once.
Persistent agent workspaces on high-IOPS block storage, orchestrated end to end with Quirq.