← Engineering blog

When It Rains Agents, Bring a mainbrella

Give an agent a file-writing tool and a command-running tool, and you can still end up with a program that can’t find the file it just wrote. The two tools may be looking at different disks. Mastra’s workspace interface lets us join them to one Mainbrella Linux container, so the agent can write a script, run it, and inspect what it produced.

We’ve added that connection in our Mastra fork. The new @mastra/mainbrella package supplies a sandbox, an HTTP filesystem, and a provider for Mastra’s editor. The examples below target that October 8 checkout.

Our Deep Agents integration followed the same broad idea: keep the agent’s existing tools and supply the machine underneath them. Mastra makes a different seam useful. Its process manager can expose a running job, its output, and its input pipe separately. That becomes interesting as soon as a command takes longer than the connection watching it.

Three mechanical assistants collaborate beneath one umbrella: one supplies tabular input, one operates a calculating machine, and one inspects its printed report. All use the same workbench.
Several workers can share one workbench. The application decides who shares it. Original AI-generated editorial illustration, rather than a captured agent run.

Two tools, one disk

Mastra is a TypeScript framework for building agent applications. A workspace gives an agent file operations and command execution without making the model learn your infrastructure API. For a small reporting task, the useful loop is familiar: inspect the input, write a program, run it, read the error, edit, repeat.

The important line in the setup below is the one that passes sandbox into the filesystem. It is the same object that runs commands. There is no second machine to keep synchronized, and no object-storage mount between the file tool and the guest.

import { Workspace } from '@mastra/core/workspace';
import {
  MainbrellaFilesystem,
  MainbrellaSandbox,
} from '@mastra/mainbrella';

const sandbox = new MainbrellaSandbox({ catalogId: 'node' });
const filesystem = new MainbrellaFilesystem({ sandbox });
const workspace = new Workspace({ sandbox, filesystem });

The defaults put both operations under /workspace. The filesystem API accepts paths relative to that root; command arguments use paths inside Linux. That distinction is small enough to overlook and large enough to ruin the report.

Two spellings for the same file, with the defaults above
From a file toolFrom a command
input.json or /input.json/workspace/input.json
report.txt or /report.txt/workspace/report.txt

With /workspace as the command’s working directory, node report.cjs sees the file written as report.cjs. If you choose another filesystem basePath, set the sandbox’s workingDirectory to match and initialize the writable filesystem before running commands there.

There’s a rough edge in this commit: stat('report.txt') returns /workspace/report.txt, but handing that value back to readFile() looks under /workspace/workspace/report.txt. The path advertised by metadata should be usable by the file tools. For now, keep filesystem calls in the workspace-relative form shown above.

The connection can leave before the job does

Suppose the reporting program has started and printed its first progress message. Then the output connection drops. Starting the program again would be a terrible recovery strategy if it had already appended a row, charged an account, or sent a message. The useful thing to recover is access to the existing job.

The process manager starts a managed Mainbrella execution and keeps its execution ID. Output events carry increasing sequence numbers. When a stream fails, the adapter can reopen it from the SDK’s last received cursor, while the original command continues. It retries eligible stream failures up to three times; it never submits a second command to repair a failed output connection.

A client starts job J once, receives output event 1, and loses its stream. Job J keeps running and records event 2 during the gap. The client reconnects with cursor 1 and receives event 2 and completion from the same job.
A hypothetical connection failure. The continuous job bar matters: the second stream watches the first execution.

In Mastra, executeCommand() uses this process manager and waits for completion. A background process gives you a handle sooner, so you can inspect output, cancel it, or send stdin. Its pid is a Mainbrella execution UUID, not an operating-system PID. An input pipe also needs an end: after feeding cat, for example, call closeStdin() so it can finish.

The result carries stdout, stderr, an exit code, and success. A timeout maps to exit code 124; cancellation maps to 137. Interruption and output-limit termination report failure too. A cheerful progress message is useful evidence that the program got somewhere. It is insufficient evidence that the report is finished.

Reconnect output

Known execution ID, last received cursor. Read the next events from that job.

Repeat a side effect

Unknown command or stdin outcome. Sending it again may do the work twice; the adapter surfaces the error for reconciliation.

Managed jobs come with a consequential limit: Mainbrella retains 32 execution records per container for one hour from admission. This adapter uses a managed job even for a short, ordinary executeCommand() call. Finishing the command releases Mastra’s local handle, but keeps the server’s record. Run 32 commands before any records expire and command 33 is rejected with execution_history_limit.

That is a real constraint on an iterative coding agent, which can spend 32 commands just getting acquainted with a repository. This version needs an execution strategy that accounts for retained history before it is a good fit for long editing sessions. Silently stopping the machine to empty the history would throw away the files we were trying to preserve.

Remember which machine did the work

There are two identities to keep straight. Mastra’s sandbox.id names the adapter in your application. Mainbrella’s remote identity contains both id and createdAt. A reusable container slot can host another machine later; the creation timestamp distinguishes this run from its replacement.

The sandbox saves that pair after startup and uses it for files, jobs, reconnects, and cleanup. Attaching to a stopped or replaced generation fails. It doesn’t quietly put the agent in an empty replacement and pretend its report is still there.

Creation needs a different recovery identity. Persist a creationKey and its creation options before starting if you need to recover across application restarts. A retry with that same key and options can reconcile an ambiguous creation within the 24-hour window. The lost-reply article follows why that distinction matters. Remembering a creation attempt and remembering a running machine solve different moments in the lifecycle.

Once several agents use a workspace, sharing becomes an application decision. A static workspace points those users at the same machine. The HTTP filesystem’s write queue orders mutations made through that filesystem instance; it cannot order a shell process’s writes or a second adapter’s changes. Its read-only flag restricts file-tool mutations, while an enabled shell can still write. Per-user isolation and permissions need to agree across both routes.

Try a small report before adding a model

The shortest useful check exercises both routes: write input through the filesystem, process it inside Linux, then read the output through the filesystem. Using the objects above, this example should print total=42:

try {
  await workspace.init();
  await filesystem.writeFile('input.json', '[12, 17, 13]');
  await filesystem.writeFile('report.cjs', `
    const fs = require('node:fs');
    const values = JSON.parse(fs.readFileSync('input.json', 'utf8'));
    const total = values.reduce((sum, value) => sum + value, 0);
    fs.writeFileSync('report.txt', 'total=' + total + '\\n');
  `);

  const result = await sandbox.executeCommand('node', ['report.cjs']);
  if (!result.success) throw new Error(result.stderr || 'Report failed');
  console.log(await filesystem.readFile('report.txt', { encoding: 'utf8' }));
} finally {
  await workspace.destroy();
}

From a checkout of the linked fork, run pnpm --filter @mastra/mainbrella build. Save the two code blocks together as workspaces/mainbrella/report-demo.mjs, or download the complete example, then run node report-demo.mjs from that directory with MAINBRELLA_API_KEY configured. The account needs an active compute allowance.

For a local Mainbrella service, set MAINBRELLA_API_URL=http://localhost:8787 and use that service’s credential. The key belongs to the application making API calls; it doesn’t need to enter the guest’s command environment.

This is an illustrative plumbing check, not a recorded model run. In reviewing the commit, we ran its 38 unit tests against the fake HTTP API, plus typecheck and lint. We also reproduced the metadata-path issue and exercised the retained-history failure with the backend’s admission limit enforced in the fake transport. The opt-in live integration suite was not run for this article.

The example reads the report before cleanup because workspace.destroy() stops this statically configured sandbox and discards its unsaved files. Saving a conversation doesn’t save that disk. Individual HTTP file transfers are limited to 1 MiB; managed output is bounded to 1 MiB, and jobs have a maximum 15-minute timeout within the machine’s lease. Export the finished artifact or explicitly save a workspace through sandbox.mainbrella if it needs to survive.

Once the file–command–file loop works, give the workspace to your Mastra agent. Then the model can meet the actual input, and you can evaluate the part no sandbox adapter can supply: whether the report is right.