Open your mainbrella with two hands
An OpenHands conversation can still know what it was doing after the address of its machine has expired. The unfinished report may be sitting on disk, with the Agent Server alive beside it. Choosing a sandbox provider means deciding how the conversation gets back there.
OpenHands is exploring more flexible sandboxing for Agent Canvas, including third-party hosting and a provider marketplace. We built a runnable Mainbrella prototype to make that discussion concrete. It connects the published OpenHands SDK to a real machine using the existing remote-workspace protocol. It also gives the proposal an awkward, useful test case: our preview endpoint can expire before the machine does.
Keep the tools. Supply their server.
OpenHands supplies the conversation and tools for a coding agent: it can edit files, run commands, and use the results to decide what to do next. Its SDK can talk to an Agent Server running elsewhere, which gives a sandbox provider a useful boundary to work with.
Our Deep Agents adapter supplies command execution and file transfers beneath the harness’s tools. Our Mastra adapter joins file tools and managed jobs to one disk. Here we can supply the whole machine hosting OpenHands’ server.
The prototype creates a Mainbrella Python machine, installs Agent Server 1.53.0 inside it, and exposes port 18000 through a protected preview. A small MainbrellaWorkspace class extends OpenHands’ RemoteWorkspace. Commands, files, conversation requests, and WebSocket events then use OpenHands’ existing protocol. Mainbrella provisions the machine; Agent Server handles the work inside it.
This matters for an iterative coding task. The agent’s file editor and terminal see the same /workspace, and its conversation talks to the server on that machine. We don’t need to recreate OpenHands’ command and event APIs in the sandbox service. The adapter source pins the SDK, server, and tools to the same version.
Provisioning has an obvious cost in this version: each new machine installs the server. A prepared image would remove that repeated setup. This example tests the connection boundary; it makes no startup-speed claim.
The credentials have different jobs too. The Mainbrella account key stays with the provisioner. The guest receives a separately generated Agent Server session key and encryption key. Model credentials follow OpenHands’ remote-conversation behavior and go to Agent Server. Moving the account key out of the guest doesn’t make the guest free of secrets.
The next connection must find the same machine
Imagine a conversation that has written half a report. An administrator changes the instance’s default provider. When the user returns, opening a fresh machine at the new default would give the agent its old conversation and an empty disk. The transcript says the report exists. The filesystem disagrees.
That is why our feedback on the proposal asks for an instance default with per-conversation and per-automation overrides, with the chosen provider/profile and runtime identity saved alongside the conversation. A default can choose where new work begins. Existing work needs its own address book.
Mainbrella’s identity includes a slot ID and a creation timestamp. A slot can be reused; the pair names the exact generation. The prototype’s attach() checks that this generation is still running. Created workspaces own cleanup; attached workspaces leave the borrowed machine running when their connection closes. Inspecting somebody’s work shouldn’t acquire a surprise delete-on-exit policy.
There is a separate identity for the attempt to create the machine. The caller supplies an idempotency key, and the adapter’s recording callback saves it before the creation request. If the reply disappears, that key lets the caller reconcile the attempt. It does not make an arbitrary shell command safe to repeat. Replaying a command that already appended a report row can give you two rows and one very confident explanation.
A machine and its endpoint have different clocks
The saved runtime handle also contains the preview endpoint, its session key, and its actual expiration. These are sensitive connection details, separate from the generation identity. In this prototype, a preview lasts at most one hour, shortened when the machine’s hard lease ends sooner.
The machine may still be running when access expires. The adapter has no automatic endpoint renewal, so reconnection needs an explicit refresh mechanism or an unavailable state. Quietly creating a replacement would conceal the problem by losing the work.
This is a good reason to make connect(handle) return credentials and their expiration. A usable provider interface has to describe how access ends as carefully as how it starts. The current preview ingress also strips application cookies. Session-key HTTP and WebSocket authentication passed the local checks; browser tooling and the embedded editor remain unqualified.
Stopping raises a different question. Ordinary stop discards unsaved guest files. Conversation history doesn’t preserve them, and a filesystem snapshot doesn’t preserve a running process’s memory. Endpoint renewal, filesystem snapshots, and process pause/resume therefore need separate capability flags. A single “supports persistence” checkbox would leave the user guessing which thing persists.
Seven checks, and a missing piece of output
The saved October 8 qualification report records seven live transport checks passing against the local Mainbrella service, using OpenHands SDK 1.53.0 and Mainbrella SDK 0.1.0. It covers authenticated server metadata; stdout, stderr, and exit status; timeout with retained partial output; a 1 MiB binary round trip; attachment to the same generation; conversation creation with a WebSocket user event; and filesystem isolation between two machines. Cleanup completed.
That list needs one particularly consequential qualification. The high-level execute_command() call reported a timeout but omitted the final buffered stdout. A second timeout probe, using start_command() and get_command_output(), received its retained partial output. The verifier recorded the high-level gap separately.
Wait for the whole command
First probe, execute_command(): timeout reported; the expected partial output was absent.
Inspect the retained command result
Second probe, start_command() / get_command_output(): the verifier received partial and exit code -1 after timeout.
The prototype README attributes this to the polling deadline coinciding with the process deadline in SDK 1.53.0. Whatever interface the maintainers choose, its qualification suite should distinguish those clocks. “The process timed out” and “the client stopped waiting” tell an agent different things about what evidence it can still recover.
The report also records seven focused lifecycle/startup tests passing from the locked environment. It records zero model requests: the WebSocket check sends a user message, without running the model. These results establish a local transport path. They do not qualify Canvas UI, hosted production, delegated workers, or the quality of an agent’s report.
What should the provider promise?
I’d start with four lifecycle operations in the SDK/Agent Server layer, with Canvas consuming that contract. These names are a proposal, not an upstream API:
| Operation | The question it answers |
|---|---|
create_or_reconcile(spec, key) | Which machine belongs to this creation attempt? |
inspect(handle) | Is that generation alive, and what limits and capabilities apply? |
connect(handle) | How do I authenticate to its Agent Server, and until when? |
stop(handle) | Did that exact generation stop, or is cleanup still uncertain? |
A provider marketplace can then show facts that affect the choice: who operates the service, what credentials it needs, actual resource and lease limits, billing constraints, network policy, and what survives stop. Providers can qualify against the same tests and report unsupported capabilities. A polished selector is useful once its options mean something inspectable.
The OpenHands issue remains an early direction-setting discussion as of October 8. This prototype works at the SDK layer; it adds no Canvas runtime option and establishes no upstream acceptance. We’ve offered it as a concrete implementation to test against the maintainers’ preferred interface.
Try the transport before adding a model
To inspect the same implementation and run its checks, start with the pinned commit:
git clone https://github.com/mainbrella/OpenHands.git
cd OpenHands
git checkout 6f3ab7771a915a5af23db6b58c49afc77e535d4a
cd examples/mainbrella
uv sync --locked
uv run --locked pytest
# Configure MAINBRELLA_API_KEY in your environment.
uv run --locked python verify.py
The README covers Python 3.12+, Node 24+ for its environment-file launcher, local setup, and credential handling. The verifier consumes two machine starts and needs allowance for two simultaneous Small machines. For a local service, use its credential and set MAINBRELLA_API_URL=http://localhost:8787; local accounts and plan access are separate from production.
Run the verifier before giving the workspace a model. You should be able to write a file, reconnect to the same disk, receive a command’s actual result, and stop the exact machine you started. If the next contract can also renew access without replacing that disk, it will answer the most useful question this prototype leaves open.