A coding agent’s memory outlives its machine
Suppose a coding agent has edited three files, found a failing test, and paused for your answer. When you say “continue,” it remembers the patch. If its temporary machine has expired, that memory describes files that no longer exist.
Our new Mainbrella provider for Open SWE has to handle that gap. It can create a Linux workspace and reconnect to it later, but it refuses to substitute a fresh machine when the old one is gone. A successful connection to an empty computer would be a particularly unhelpful kind of success.
The October 8 commit adds the provider to our fork. Its interesting parts are the promises around a resumed job—and two places where reviewing the surrounding code revealed that the adapter hasn’t fulfilled the whole contract yet.


The workspace belongs to a conversation
Open SWE turns an engineering request into a longer-running job: investigate a repository, change code, validate it, deliver a pull request, and respond to feedback. Requests can arrive through GitHub, Slack, or its dashboard. Deep Agents supplies the planning and file tools; LangGraph carries execution and conversation state; Open SWE adds the engineering workflow around them.
That means the shell is more than a calculator attached to one model call. A follow-up in the same thread can depend on an uncommitted edit, installed packages, and the checkout that produced yesterday’s test failure. Open SWE records a sandbox ID in the thread’s metadata and uses it to reconnect. Independent conversations can have independent workspaces.
Mainbrella joins the existing provider registry. The new adapter implements asynchronous command execution, binary uploads and downloads, and sandbox identity over HTTP, using the project’s existing httpx2 dependency. The registry awaits it directly. There’s no additional provider SDK to install, and the existing higher-level file tools can keep using their familiar backend interface.
It’s a small seam in the application, but it carries a substantial assumption: every tool call is addressed to the machine containing this thread’s work. The thread lifecycle already treats an unreachable working sandbox as an error. The Mainbrella adapter preserves that caution.
A slot name can stay while its contents disappear
Consider our paused patch in container slot c1. The machine expires, then another job starts in the same slot. Looking up “c1” will find a running computer. It won’t find the computer that ran our test.
The adapter’s identity therefore includes the slot and the exact creation timestamp. An illustrative ID looks like this:
mainbrella:c1@2026-10-08T12:00:00.000Z
The timestamp identifies a generation: this particular incarnation of the slot. Reconnection checks the account’s container list for the same slot, timestamp, and running status. Commands and file transfers also send both values, so a machine that expires between the initial check and a later write can’t redirect that write into its replacement.
Here’s the choice that matters to the person waiting for the patch:
Accept any running c1
The agent resumes with the old test result and a different filesystem. It can mistake missing work for a new fact about the repository—or start editing another job’s workspace.
Require the recorded generation
The adapter raises when that generation is missing, stopped, or replaced. Recovery can address the missing work explicitly, instead of presenting a replacement as continuity.
This provider doesn’t automatically save or restore Mainbrella workspaces. Open SWE’s LangSmith snapshot settings don’t apply to it either. Git commits pushed to the remote and files exported before expiration can survive; unexported guest files are discarded when the machine stops. Recording a more precise address prevents the wrong attachment. It doesn’t preserve the disk.
The retry depends on what disappeared
The machine can also be alive while the HTTP reply is missing. During creation, the adapter generates one idempotency key and keeps using it for retries within a 120-second attempt. It retries transport failures and HTTP 503 responses with the same image and size selection. The backend can return the original creation rather than charge another start for a duplicate launch.
It doesn’t simply pick a running container from the account list. The creation receipt names the slot and generation it reserved; the adapter waits for that exact creation to be running and confirmed in the list. If the attempt remains ambiguous at its deadline, the exception includes the creation key for reconciliation. That key belongs to this attempt; starting the factory again generates a new one.
Now replace “create a machine” with “append a line to a file.” The same automatic retry could append the line twice. Foreground commands and file writes aren’t automatically replayed after a lost response. A model or application deciding what to do next still needs to distinguish a failed command from a command whose outcome it didn’t receive.
For commands with longer timeouts, the adapter starts a managed job and polls its retained result. Once it has the job ID, cancellation of the wait requests cancellation of that job before propagating the interruption. That cleanup starts after admission: cancellation while the initial POST is in flight isn’t covered by the cleanup block. The backend keeps jobs running when a client disconnects, so interrupting the agent at that point can leave work running until its own deadline.
These are different recovery problems. Retrying creation can find the reserved machine. Reading a known job can find its result. Repeating arbitrary shell text can perform the work again. A blanket “retry on network error” would erase those distinctions.
The 60-second default that became five minutes
The adapter uses foreground execution for timeouts up to 60 seconds. Longer timeouts, up to 900 seconds, use managed jobs. That choice matters because Mainbrella retains at most 32 managed-job records per generation for one hour from admission, including jobs that have finished. The 33rd admission before a record expires returns 429 execution_history_limit.
Reading the adapter alone suggests that an ordinary git status won’t consume one of those records: its own default is 60 seconds. Reading the wrapper changed that conclusion. Open SWE’s per-thread proxy supplies 300 seconds when the caller omits a timeout. The adapter receives an explicit five-minute allowance and chooses managed execution even if the command finishes immediately.
| Caller | Timeout received | Execution |
|---|---|---|
| Adapter directly; timeout omitted | 60 seconds | Foreground; no retained job |
| Open SWE thread proxy; timeout omitted | 300 seconds | Managed; retains a job record |
| Thread proxy; explicit 60 seconds | 60 seconds | Foreground; no retained job |
A local probe through the actual proxy confirmed that routing. This isn’t a measured production failure rate; it’s a deterministic mismatch between the wrapper’s default and the provider’s admission limit. A busy coding session can exhaust its managed history with short shell calls. The integration needs a deliberate timeout policy before that path is suitable for sustained work. Automatically retrying a timed-out foreground command as a managed job would introduce the replay problem above.
There’s a second inheritance trap. Most asynchronous file helpers in the installed Deep Agents BaseSandbox call the adapter’s async primitives, but deletion falls back to running the synchronous delete method in a worker thread. That method calls execute, which this async-only adapter deliberately rejects. Calling adelete reproduced a NotImplementedError before any HTTP request. File deletion through that helper needs its own async implementation; shell deletion remains available.
The nine new provider tests pass. They cover creation retries, exact-generation reconnection, foreground results, managed polling and cancellation, and partial-success binary transfers against mocked HTTP responses. They don’t exercise the thread proxy’s default or inherited deletion. Those two findings are a useful reason to test an adapter through its caller as well as through its own methods.
A healthy machine still needs the right tools
The default catalog image is Python because the inherited file tools run python3 helpers inside the guest. A Node project still needs Python for those tools. Custom images should include Python, bash, git, and GNU coreutils, plus the project’s runtimes and Open SWE’s expected tools such as gh and rg. A workspace setup script can supply missing tools; the Python catalog image by itself isn’t a complete coding-agent environment.
The adapter also starts commands in /workspace and restores HOME from the guest user’s account when it’s absent. That detail serves a concrete caller: Open SWE reapplies global git identity on resumed runs, and Mainbrella’s execution environment inherits only PATH. “Git is installed” is weaker than “the agent can configure git the way its workflow expects.”
Other boundaries remain visible in the result: binary transfers and command output are capped at 1 MiB; a timeout preserves partial output and reports exit code 124; truncation retains its flag. A command’s timeout never extends the machine’s hard deadline. The adapter waits for command results; it doesn’t expose a live terminal or automatically add the platform’s other APIs as agent tools.
Try the boundary before entrusting it with a patch
Use the fork and commit linked above, with the existing Open SWE development setup. The base uv sync install includes what this provider needs. Add these values to Open SWE’s environment, using a key from the Mainbrella deployment you intend to call:
SANDBOX_TYPE="mainbrella"
MAINBRELLA_API_KEY="mb_..."
MAINBRELLA_CATALOG_ID="python"
MAINBRELLA_SANDBOX_SIZE="small"
An optional MAINBRELLA_IMAGE_ID selects a custom image instead of the catalog. MAINBRELLA_API_URL defaults to the hosted HTTPS API and accepts loopback HTTP for a local backend. Authentication and compute allowance are separate: the target deployment must also grant active access. The customization guide records the provider’s configuration, and the API contract describes retained-job limits.
The checks behind this article were local mock-contract tests and the two focused caller probes, not a live Open SWE agent run or a hosted performance benchmark. The timeout routing and deletion findings remain open in the reviewed commit.
Before asking an agent to resume valuable work, decide where that work survives a stopped machine. A pushed branch or an explicitly saved and restored workspace can supply an answer. A conversation that remembers “I changed three files” can’t.