Benchmarks

Measured startup, dependency installation, and build/test times across Lite–XL, plus historical prototype startup percentiles. These observations are not latency or availability guarantees.

Five-size workload measurements

October 6, 2026 UTC. Each cell below is a single observed public-API phase time in seconds, including managed execution polling. Workload directories and package caches were fresh; provider host and image-cache placement were unknown. These samples do not establish P50 or P95.

SizeNode startupnpm installTypeScript build/testpip installPython build/numeric test
Lite1.42 s4.36 s4.49 s9.38 sInvalid fixture
Small1.04 s4.20 s4.09 s9.03 s2.69 s
Medium0.96 s4.01 s2.51 s9.20 s2.40 s
Large0.88 s4.10 s2.31 s9.32 s2.50 s
XL0.90 s4.03 s2.58 s7.47 s2.34 s

The Node workload installs TypeScript 5.9.3 and esbuild 0.25.11, then compiles, bundles, and verifies 200 generated modules. The Python workload installs requests 2.32.5 and NumPy 2.3.4, compiles 200 modules, and checks a 200 × 200 matrix calculation.

Two simultaneous Small Node generations started in 0.90 and 1.19 seconds, then completed combined npm installation/build/test in 7.49 and 7.56 seconds. This is a two-generation observation, not a concurrency saturation result.

The batch requested 17 customer starts; all recorded generations were cleaned up. Six complete generation workflows succeeded and eleven failed; successful Node/Python phases from workflows with later failures are included above. Browser installer output exceeded the managed-output limit on Small–XL and timed out on Lite. Rust environment failures and a Lite Python fixture bug invalidate those workload results. Corrected fixtures still require live replay.

Download timing samples and run metadata (identifiers, credentials, and command output omitted). The qualification record was added in backend commit 8c79fd5; recorded runner revision: 2323129. See the size runner and corrected workload definitions.

Historical prototype startup

October 5, 2026, 11:43–11:53 UTC. A separate Cloudflare-backed prototype used standard-2: 1 vCPU, 6 GiB RAM, and 12 GB disk. Its startup protocol differs from the current authenticated API; these timings are not directly comparable with the five-size workload run.

TestAttempted / successfulP50P95
Create30 / 30381 ms528 ms
Snapshot restore30 / 30473 ms788 ms
100 concurrent starts100 / 100503 ms680 ms

The original harness and report were recovered from the prototype reference files. P50/P95 use nearest rank over successful client start-request wall times, including response decoding. A separate shell check confirmed usability after each start; that check is excluded from the reported startup latency.

The create test ran sequentially. Snapshot restores used 30 new identities and verified a marker plus the SHA-256 of a 30 MiB random asset alongside 300 source files. The burst issued 100 starts within 4.54 ms, confirmed unique hostnames, and observed all 100 running simultaneously. All 164 machines created across the full suite were confirmed stopped.

Fresh identities do not prove cold hosts or uncached images. Client location, compute region, and exact image digest were not recorded. Snapshot restoration here describes the prototype experiment, not a customer feature availability claim. Download timing samples · Original prototype harness.

Reproducible HTTP methodology

The runner below measures the current authenticated API, using the same lifecycle, foreground execution, binary files, managed streaming and cancellation, and generation-qualified cleanup as the Quickstart. This startup runner is a separate protocol from both the five-size workload runner and the historical prototype harness.

Create latency
Client wall-clock time from the first creation POST to a confirmed running response, including response decoding, retries, and polling with the same idempotency key.
First-command latency
Time from that first POST to a validated hello command response. Includes creation and the HTTP execution round trip; excludes preflight, file verification, managed execution/cancellation checks, and cleanup.
Samples and percentiles
All attempted samples remain in the JSON report, including errors and cleanup failures. P50 and P95 use nearest rank over fully successful samples only. Requested and attempted counts are separate; unsuccessful samples never become zero-millisecond results.
Concurrency and cache state
Runs in batches of up to the chosen concurrency, with new reservations and no prewarming by the runner. Request arrival, host reuse, and image caches are uncontrolled. A concurrent HTTP batch does not prove all machines overlapped in their running state.
Environment
The report includes UTC run times, Node version, API origin, selected catalog image, plan, returned limits, and a client location you supply. Compute region and provider cache state remain unknown because the API does not report them.

Run and retain the evidence

Use Node 22+, an active account, and MAINBRELLA_API_KEY from your secret manager. Each attempted reservation consumes one monthly start, including failures. The runner checks allowance for the entire requested run and free slots for a batch. Avoid other launches during the test. Resolve any ambiguous creation or failed cleanup before running again.

curl --fail --silent --show-error https://mainbrella.com/mainbrella-doctor.mjs -o mainbrella-doctor.mjs
curl --fail --silent --show-error https://mainbrella.com/mainbrella-verify.mjs -o mainbrella-verify.mjs
curl --fail --silent --show-error https://mainbrella.com/mainbrella-benchmark.mjs -o mainbrella-benchmark.mjs
# Inspect all three scripts before running.
node mainbrella-doctor.mjs
MAINBRELLA_CLIENT_LOCATION='your city / cloud region' \
  node mainbrella-benchmark.mjs --samples 30 --concurrency 1 > benchmark.json
# For an account with at least 100 free slots and starts:
# node mainbrella-benchmark.mjs --samples 100 --concurrency 100 > burst.json

Download the runner · Source code. Preserve the runner version with the JSON report. Review container identifiers before sharing reports. Publish raw samples, client location, runtime image, failures, and cleanup outcomes alongside any new performance claim.