“Resume takes one second” is incomplete unless the timer has a defined finish line. It could mean that an API accepted a request, that a process started, or that an agent successfully continued its job. Those events can be separated by scheduling, image loading, storage attachment, checkpoint replay and network reconnection.
TL;DR
- Time resume from the request to the first correct application result.
- Record API or SDK return, execution readiness and useful work separately; a fast acknowledgement is not ready-to-run recovery.
- Keep resource sizes, retained state, cache conditions, concurrency and readiness probes comparable.
- Report failed recoveries and timeouts alongside latency; calculate end-to-end percentiles from complete trials.
Measure from the controller's request to the first correct application result, and retain the intermediate milestones. That gives you a user-relevant latency and enough detail to explain where the time went.
Methodology updated 7 October 2026. This article provides a test design and illustrative harness; it does not report new provider measurements.
Where should the timer start and stop?
Use one controller-side monotonic clock for the end-to-end measurement. Python's perf_counter_ns() provides a high-resolution duration clock. Do not subtract timestamps from different machines unless their clock relationship has been established.
| Milestone | Definition | What it tells you |
|---|---|---|
| Request start | Immediately before invoking the resume operation | The beginning of the user's wait |
| API or SDK return | The resume call returns successfully | What this particular call guarantees |
| Execution reachable | A small command completes inside the recovered environment | The execution path is available |
| Useful work succeeds | A state-dependent application probe returns the correct result | The agent can actually continue |
Name the second milestone honestly. Some SDKs wait for readiness before returning; others expose an asynchronous operation. A blocking SDK return is not a pure API-acknowledgement time. Record the SDK version and behavior so another engineer can reproduce your boundary.
Nirvana's pause/resume guide says that resume starts pod creation and documents a Pending state. That is why an “instant” lifecycle response should not be presented as a measurement of ready-to-run recovery.
Use separate timers for pause and compute release. A pause that retains files and a pause that captures RAM perform different work. Neither should be compared with destructive teardown as if all three were the same operation.
What should the application probe actually do?
An empty echo verifies transport and command execution. It does not prove the application loaded the correct checkpoint, mounted the right files or reconnected to its database.
For a coding agent, the final probe could verify a repository revision and an uncommitted file hash, load the saved task ID, then run a small deterministic test. For a data agent, it could read a saved cursor and produce the next expected output from a fixed input. Keep the probe small enough to isolate recovery while still exercising the state your users depend on.
Use a fresh random memory token before pausing if process continuity is a requirement. Checking a PID alone is weaker because identifiers can be reused. Keep an external expected value so the restored environment cannot accidentally validate itself against newly initialized state.
Reconnect remote dependencies explicitly. E2B's persistence documentation notes that clients must reconnect after pause. A restored process and a healthy connection to an outside service are distinct outcomes.
The following asynchronous skeleton uses Python 3.11 or later and shows the measurement boundaries. The three callables are adapters you implement for your SDK and application; this is not drop-in provider code. Each adapter also needs appropriate transport timeouts and cancellation behavior.
import asyncio
from time import perf_counter_ns
async def measure_resume(resume, exec_probe, work_probe, timeout_s=60):
started = perf_counter_ns()
marks = {}
result = {"ok": False, "timeout_s": timeout_s}
def mark(name):
marks[name] = (perf_counter_ns() - started) / 1_000_000
try:
async with asyncio.timeout(timeout_s):
await resume()
mark("call_return_ms")
await exec_probe()
mark("execution_ready_ms")
await work_probe() # raises if recovered state is incorrect
mark("useful_work_ms")
result["ok"] = True
except Exception as exc:
result["error_type"] = type(exc).__name__
finally:
result["elapsed_ms"] = (
perf_counter_ns() - started
) / 1_000_000
result.update(marks)
return result
The probes should retry only the explicitly expected “not ready yet” condition, within the deadline. Authentication failures and incorrect recovered data should fail the trial. Otherwise a permissive retry loop can hide a broken system behind a slow success.
Which conditions need separate experiments?
Record a scenario alongside every result. At minimum include provider, region, environment class, CPU and RAM, image digest, persistence mode, SDK version, controller location, workspace size and file count.
Separate a new environment, a reused warm environment and a restored user workspace. GKE's warm-pool guide describes assigning an available pod to a claim. That operation is not the same as rebuilding an accumulated workspace from scratch. Daytona also documents warm pools; identify whether your test uses one.
Use a small experiment matrix rather than changing everything at once:
| Dimension | Example scenarios | Why it matters |
|---|---|---|
| Retained state | Empty workspace; representative repository; large accumulated workspace | Restore and initialization work can differ |
| Memory | Small process; realistic loaded application | Relevant when the selected operation captures RAM |
| Cache condition | Known warm image; confirmed uncached image | Image preparation can dominate a fresh start |
| Load | One recovery; a defined burst of simultaneous recoveries | Queueing and capacity become visible |
| Idle duration | Brief pause; normal overnight gap | Exposes retention and reconnect behavior |
Keep the starting fixture reproducible. Record what changed if a provider requires a different image or resource size. Interleave comparable trials across providers to reduce the chance that one system is measured only during a quiet period.
How should failures and tail latency be reported?
Publish the number of attempts, successful recoveries, timeouts, incorrect-state failures and other errors. Then report latency percentiles for successful runs, clearly identifying that denominator. A timeout is not a fast result and should not disappear from the reliability calculation.
Small samples make extreme percentiles unstable. With 100 successful runs, a p99 estimate is determined by observations near the maximum; it is not a precise forecast of production tail behavior. Use exploratory runs to find problems, then collect enough representative observations for the claim you intend to make.
Do not add p95 values for individual phases and call the sum an end-to-end p95. The slowest API calls and the slowest application recoveries may occur in different trials. Calculate the total for each trial first, then take the percentile of those totals.
Record polling intervals too. A readiness check every second can add almost a second of detection delay. Keep polling behavior consistent, or disclose the difference when the goal is to measure each provider's normal client experience.
What evidence makes the result useful to another team?
A reproducible report includes the harness revision, fixture manifest, raw per-trial records, failure categories and the exact readiness probe. Retain a trace from at least one normal recovery and one slow or failed recovery. A summary number without this evidence is difficult to act on.
Report the user-facing outcome first: “the saved task resumed correctly within the application budget in X of Y attempts.” Fill those values only from actual measurements. Follow with the phase timings that explain the outcome and the configuration that produced it.
For a Nirvana evaluation, use the documented lifecycle and test checkpoint loading in the final probe. The result should tell an engineer whether their agent can continue, not simply whether a control-plane request returned.
FAQ
Should pause time be included in resume latency?
Report it separately. If the product experience includes both operations, also report the complete cycle with its own start and finish boundaries.
Is a provider's “Running” status sufficient?
It is a useful milestone. The final acceptance check should still exercise the state and dependency needed by the application.
Can a faster API response still produce slower recovery?
Yes. Work after acknowledgement can dominate. That is the reason to retain phase timings and the end-to-end application result.
About Nirvana Labs
Nirvana Labs is a high-performance storage cloud purpose built for blockchain, AI and databases i.e. the most demanding, real-time, stateful workloads. Accelerated Block Storage (ABS) offers 20K baseline IOPS included, no over provisioning. Nirvana Kubernetes Service (NKS) with Karpenter auto-scaling, high clock-speed compute and private networking. Backed by Jump Trading, Crucible, etc with 50+ customers live in production today.
Learn more at Nirvana Labs
Nirvana Cloud | Pricing | Blog | Docs | Changelog | LinkedIn | Twitter | Telegram | YouTube
Related Posts

How to Benchmark PostgreSQL WAL and Commit Performance on Cloud Storage
A repeatable PostgreSQL storage evaluation: control durability settings, separate transaction and commit latency, inspect WAL and checkpoint metrics, and test recovery.

How to Test Redis AOF Rewrites and Recovery Before Production
A practical acceptance test for Redis persistence: measure rewrites under load, validate acknowledged writes after failure, and prove independent backup recovery.

Altinity.Cloud BYOC: An Evaluation and Migration Guide
Evaluate Altinity.Cloud BYOC for ClickHouse® with ownership, workload, backup and migration checks before a reversible cutover.

