← Engineering blog

Fifteen companies, 480 containers, one awkward profit question

Margin of Error Ltd had a problem: sales were growing, but profit was shrinking. Someone needed to explain where the money went. We gave the problem to a collection of Linux containers. Then we gave fourteen other companies the same treatment.

The companies were synthetic. Their invoices, returns, inventory, and suspiciously generous discounts were synthetic too. The compute was real: fifteen private service networks, thirty-two containers per network, and a report pipeline that had to finish with numbers we could check. A fictional business can still have a very real double-counting bug.

We built this as a substantial Mainbrella demo. A viewer could ask a company a question, watch its analysts work, inspect a failed attempt, and open the resulting HTML report and CSV files. The interesting part was getting all those operations to agree about whose data they were processing, which attempt owned the work, and whether the final answer was actually supported by the records.

Conceptual cutaway of three separate company networks. Each contains shared service modules, worker modules, and its own report output under a blue umbrella.
Each company had its own service namespace. The illustration shows three companies; the complete layout used fifteen networks with thirty-two members apiece.

What counted as a company

A company was one network and a small reporting application distributed across its members. Four containers supplied shared services; the other twenty-eight ran analysis jobs. We kept the code in Node, used small JSON shards, and gave every role a useful job.

One company: 4 service containers + 28 workers = 32 members.
RoleCountResponsibility
Coordinator1Task leases, attempt IDs, progress, and completion records.
Data service1Immutable input shards, schemas, and content hashes.
Artifact service1Accepted results, CSV exports, and the rendered report.
Validator1Independent checks on totals, coverage, and provenance.
Analysts28Partitioned calculations for revenue, returns, costs, and margin changes.

The coordinator handed out bounded tasks: compare two months of returns in a region, compute the effect of a discount, reconcile a product category. Those tasks were easy to partition and cheap enough for Lite machines, which have 256 MiB of memory. A container could answer one question without loading an entire company’s history.

The LLM-facing controller lived outside the company containers. It turned questions into analysis plans and, when needed, generated analysis code. Model credentials and the Mainbrella API key stayed with that controller. The workers received code and task data. The validator did arithmetic.

The launch manifest was part of the application

Before launching anything, we read GET /capabilities and authenticated GET /containers. The deployment had to advertise private services, the selected image, and the execution and file operations we needed. The account needed enough starts, container slots, concurrent compute units, and runtime allowance for the batch.

Our Scale account allowed 500 simultaneous containers and 640 concurrent compute units. The full layout used 480 Lite machines, each costing one concurrent unit. The network contract allowed sixteen networks per account and thirty-two members per network. Fifteen companies fitted those limits. The leftover network slot was useful breathing room, rather than an invitation to found another conglomerate.

We recorded every intended creation before sending its request. This is the core of the server-side launch code; a batch runner called it with bounded concurrency.

// Server-side Node. Keep this key out of company containers.
import { randomUUID } from 'node:crypto';
import { mkdir, writeFile, rename } from 'node:fs/promises';

const origin = process.env.MAINBRELLA_API_URL
  ?? 'http://localhost:8787';
const key = process.env.MAINBRELLA_LOCAL_API_KEY
  ?? process.env.MAINBRELLA_API_KEY;
if (!key) throw new Error('Provision a server-side API key');

async function api(path, { method = 'GET', body, headers = {} } = {}) {
  const response = await fetch(new URL(path, origin), {
    method,
    headers: {
      Authorization: `Bearer ${key}`,
      ...(body === undefined ? {} : { 'Content-Type': 'application/json' }),
      ...headers,
    },
    body: body === undefined ? undefined : JSON.stringify(body),
    redirect: 'error',
    signal: AbortSignal.timeout(90_000),
  });
  if (!response.ok) throw new Error(`Mainbrella HTTP ${response.status}`);
  return response.json();
}

async function journal(name, value) {
  await mkdir('.mainbrella', { recursive: true, mode: 0o700 });
  const path = `.mainbrella/${name}.json`;
  await writeFile(`${path}.tmp`, JSON.stringify(value), { mode: 0o600 });
  await rename(`${path}.tmp`, path);
}

const operation = {
  key: randomUUID(),
  issuedAt: new Date().toISOString(),
  body: { catalogId: 'node', size: 'lite' },
};
await journal(operation.key, operation); // Before the network request.
const started = await api('/containers', {
  method: 'POST', body: operation.body,
  headers: { 'Idempotency-Key': operation.key },
});
const { creation } = started;
if (creation?.status !== 'running') {
  throw new Error('Reconcile the saved creation key and body');
}
const machine = started.containers.find(c =>
  c.id === creation.containerId && c.createdAt === creation.createdAt
  && c.status === 'running');
if (!machine) throw new Error('Running generation missing');
const identity = { id: machine.id, createdAt: machine.createdAt };
await journal(operation.key, { ...operation, identity });

That little journal matters when the response disappears. The backend might have reserved a slot and started a machine already. We reconciled an ambiguous creation with the same body and Idempotency-Key, within the twenty-four-hour recovery window. A fresh UUID would describe a fresh launch.

The implementation is readable in container-account-core.js. The account Durable Object serializes admission, reserves the slot and quota, and stores the keyed creation record. Boot happens outside the reservation lock. Pending slots still count toward capacity, so concurrent requests cannot all claim the same last available machine.

We kept both id and createdAt for every running container. IDs identify slots; the returned timestamp identifies a particular generation in that slot. An analyst called c17 today can be a different analyst after a stop. A delayed cleanup request needs enough information to tell them apart.

A worker asked for a task at coordinator.internal

Each network reused the same service names: coordinator, data, artifacts, and validator. We registered a service’s application port when it accepted requests. Workers joined as caller-only members, without an exposed service port.

// coordinator and worker are previously recorded { id, createdAt } pairs.
const network = 'margin-of-error';
await api('/private-services/networks', {
  method: 'POST', body: { name: network },
});
const membership = '/private-services/members?'
  + new URLSearchParams({ network });
await api(membership, {
  method: 'PUT',
  body: { ...coordinator, name: 'coordinator', port: 8080 },
});
await api(membership, {
  method: 'PUT',
  body: { ...worker, name: 'analyst-01' }, // Caller-only member.
});

// Inside that worker: port 80 is the private HTTP routing entry point.
const response = await fetch('http://coordinator.internal/tasks/next');
if (!response.ok) throw new Error(`Queue HTTP ${response.status}`);
const task = await response.json();

The URL inside the worker used plain HTTP on port 80. The registered coordinator listened on port 8080. Mainbrella resolved that indirection using the source’s account, network membership, and exact container generation.

private-services-runtime.js installs an outbound HTTP interceptor for *.internal. Its relay strips reserved identity headers from the guest request and supplies trusted source identity through the runtime entrypoint. The worker cannot select another account by inventing an x-mainbrella-user header.

In PrivateServicesController.route(), the account-owned registry finds the source network and a destination member with the requested name and a registered port. The destination rechecks its generation and registration under its lifecycle lock before opening the application port. The routing path also rechecks membership around response release, because a detach or replacement can race an in-flight request.

This was private HTTP service routing. It did not turn the demo into a general TCP network, provide PostgreSQL connections, or change outbound internet policy. The prototype bounded request and response bodies to 1 MiB and requests to ten seconds. We served small shards and paginated records. Every company could call its own data.internal; a report from Margin of Error could not resolve a differently named service registered only in another company’s network.

Private Services required networking.privateServices: true and deployment enablement. This project exercised the local prototype. Those routing checks describe the feature we used; they do not establish an enterprise tenancy or compliance claim.

A task lease was a different thing from a container lease

We still needed an application scheduler. Mainbrella supplied compute and command lifecycle; the coordinator supplied business tasks. Its queue leased a shard to a particular attempt and accepted one valid completion for the task.

// Application protocol; Mainbrella does not supply this task queue.
{
  "runId": "profit-2026-09",
  "taskId": "returns-west",
  "attemptId": "returns-west-02",
  "leaseToken": "<opaque token issued by the coordinator>",
  "inputSha256": "<hash of the assigned input shard>",
  "leaseExpiresAt": "<application task deadline>",
  "inputUrl": "http://data.internal/shards/returns-west.json"
}

A worker could die after calculating a result but before acknowledging completion. The coordinator could reassign the task after its application lease expired. An old attempt might then wake up and submit its answer. We checked the attempt ID, lease token, deadline, and input hash before accepting that answer. This kept a late worker from overwriting a newer result.

Completion was deduplicated by (runId, taskId) in the application’s durable records. Input shards were immutable. Results carried the shard hash and analysis version. We could inspect what a worker had calculated without guessing which dataset happened to be current when it started.

A task’s deadline described ownership of a piece of analysis. The container’s lease described how long the compute could exist. Expiring one did not magically renew the other.

Longer jobs got a retained execution ID

For a quick probe, foreground POST /containers/exec was enough. Its maximum timeout was sixty seconds. For an analyst processing a batch or a shared service running for a demonstration segment, we used managed execution and retained the job ID alongside its container identity.

// worker.mjs has already been uploaded into /workspace/demo.
const query = new URLSearchParams(worker);
const execution = {
  key: randomUUID(),
  body: {
    argv: ['node', '/workspace/demo/worker.mjs'],
    timeoutMs: 120_000,
  },
};
await journal(execution.key, { ...execution, container: worker });
const job = await api(`/containers/executions?${query}`, {
  method: 'POST', body: execution.body,
  headers: { 'Idempotency-Key': execution.key },
});
await journal(execution.key, { ...execution, container: worker, jobId: job.id });

// SSE: /containers/executions/{job.id}/events?...&cursor=0
// Poll: /containers/executions/{job.id}?...
// Cancel: DELETE /containers/executions/{job.id}?...
// Closing the SSE connection only detaches the viewer.

The argv form passed literal arguments without constructing a shell command. The two-minute timeout fitted under managed execution’s fifteen-minute maximum; every job was also bounded by its container’s remaining lease. Shared services were restarted deliberately between demonstration segments as their managed jobs ended. A permanently running company would need an explicit service lifecycle beyond that segment.

executions.js persists the execution identity before starting the process. Matching keyed retries recover the same job for the retained one-hour window. SSE events carry increasing sequence numbers. On reconnect, the controller resumed with the last processed sequence as cursor and ignored duplicates.

Closing a browser tab detached its stream. Cancellation required deleting the retained execution and checking the resulting terminal status. We also checked exitCode, timedOut, and outputTruncated; a clipped JSON result with a cheerful opening brace was not a completed analysis.

The runtime allowed four concurrent command/file operations per container, with bounded output. Managed history retained up to thirty-two jobs during its retention window. We used one analyst job to process a finite batch of task leases, rather than generating a separate retained execution for every invoice.

Data sheets feed parallel worker modules, then a validation stage. Accepted results flow into a report and spreadsheet; discrepancies branch into a separate exception tray.
The report pipeline: immutable inputs → parallel analysis → independent validation → artifacts. Exceptions remained visible instead of being blended into the final totals.

The accountant did not get a temperature setting

The report could contain an LLM-written explanation. Its financial totals came from deterministic code. We represented money in integer cents, validated schemas, and retained row counts and input hashes. For larger totals, the same design could use decimal arithmetic or BigInt with an explicit serialization format.

// Money stays in integer cents. Inputs are schema-checked first.
const profitCents = revenueCents - cogsCents - operatingExpenseCents;
const explainedChangeCents = contributions.reduce(
  (sum, item) => sum + item.deltaCents, 0,
);
assert.equal(explainedChangeCents, currentProfitCents - previousProfitCents);
assert.equal(result.inputSha256, assignedTask.inputSha256);
assert.equal(result.attemptId, assignedTask.attemptId);
assert.equal(result.leaseToken, assignedTask.leaseToken);
assert.ok(Date.now() < Date.parse(assignedTask.leaseExpiresAt));
// Atomically accept at most one valid completion for this runId/taskId.

The validator independently recomputed the relevant aggregates and checked that the contributions explained the profit change. Missing shards, duplicate completions, mixed analysis versions, and expired attempts were errors. A worker’s exit code of zero established that its process had finished successfully. It did not establish that subtraction had occurred in the correct direction.

For each company, the artifact service assembled report.html, metrics.csv, exceptions.csv, and a manifest containing input hashes, analysis versions, and accepted task IDs. The narrative linked to the calculations behind its claims. Every reported cause had an inspectable trail back to the source records.

We escaped data when rendering HTML, kept generated scripts outside the report, and rendered the report without executing arbitrary worker-supplied markup. A company name was a string, even when it had strong opinions about closing a <script> tag.

The controller retrieved bounded files through GET /containers/files, with encoded paths and the exact generation. File writes used raw bytes through PUT, required an existing parent directory, and replaced regular files atomically. The write path in files.js rejects symlink targets; reads can follow symlinks inside the owned guest. Larger exports needed an explicit chunking or portable-export workflow. None of these files belonged in stdout.

The 254th local address was already spoken for

The first infrastructure problem arrived during provisioning. After 253 task containers had attached, three launch requests returned HTTP 503. The account allowance was large enough. Docker’s local default bridge was a /24, and its 253 usable container addresses were occupied.

That was a local runtime networking limit, separate from Mainbrella’s fifteen logical company networks. We paused launches, retained the ambiguous creation keys, and moved only task-owned Docker proxies to a larger, dedicated bridge. We identified each proxy by running hostname through the recorded container generation, then checked command execution after moving it. All 253 migrated proxies passed that check.

The three failed creations were subsequently confirmed stopped through their original keys. Only then did we prepare intentional replacement launches. Provisioning resumed and the final check matched all 480 running generations against fifteen membership lists of thirty-two.

That detour belonged in the demo. A dropped launch response should leave a recoverable creation record, rather than a 481st employee nobody remembered hiring.

Saving the work came before stopping the machine

Stopping ordinarily discarded the container filesystem. We used saved workspaces when we wanted a source environment available for later restoration, and copied report artifacts outside the container before teardown.

const save = {
  key: randomUUID(),
  body: { ...worker, name: 'margin-of-error-analyst-01', stop: false },
};
await journal(save.key, save);
const workspace = await api('/workspaces', {
  method: 'POST', body: save.body,
  headers: { 'Idempotency-Key': save.key },
});
if (workspace.status !== 'ready'
    || workspace.source.id !== worker.id
    || workspace.source.createdAt !== worker.createdAt) {
  throw new Error('Save unresolved; preserve the source container');
}
await journal(save.key, { ...save, workspaceId: workspace.id });
const after = await api('/containers?' + new URLSearchParams(worker), {
  method: 'DELETE',
});
if (after.containers.some(c =>
  c.id === worker.id && c.createdAt === worker.createdAt)) {
  throw new Error('Stop unresolved; reconcile this exact generation');
}

The backend’s workspaces.js reserves save capacity before capture, persists the provider handle, and optionally stops the source only after committing that snapshot. The snippet used separate save and stop requests so the controller could durably record the workspace ID between them. The API also supports stop: true in the save body.

Our account could retain one hundred saved workspaces. We saved the first hundred containers in manifest order, confirmed each workspace was ready for the exact source generation, and then confirmed those generations were absent from the running list. The remaining 380 containers stayed running. We did not delete saved workspaces to make a success counter look tidier.

A restore used POST /containers with { workspaceId } and a fresh persisted creation key. It consumed a start and returned a new generation. Filesystem state survived; RAM, processes, previews, and network memberships did not. Restored services needed to start again and register their new identities.

The reporting application therefore kept progress in files or durable coordinator records and quiesced writers before taking a meaningful checkpoint. An in-memory queue would vanish on restore. The one hundred infrastructure saves proved the save/stop path; their count alone did not prove that a completed report could resume halfway through rendering.

What we put in front of the viewer

The demonstration began with a company and a question. It showed leased tasks, accepted results, validation failures, and links to the generated artifacts. Selecting a task opened its container generation, managed execution ID, logs, and input hash. Canceling a slow attempt left the coordinator enough information to reassign its work.

We kept provisioning measurements separate from dispatch into an already-running worker. We also kept useful throughput separate from raw process count: completed, validated tasks per second; error rate; command latency; time until the report was available; and whether the artifact checks passed. Four hundred and eighty running containers described capacity. The files at the end described what that capacity accomplished.

Everything had a deadline. Status polling did not renew idle time. Managed execution could keep a container active while its job ran, but the hard deadline still applied. Saved workspaces had returned expiration dates and finite retention. The controller’s final responsibility was to export outputs and stop only the generations it owned.

For the implementation trail, start with the open-source Mainbrella backend, the HTTP contract, and the agent setup workflow. Follow a creation key through admission, a generation through a private request, and a task ID into the artifact manifest. Then open metrics.csv and check the profit calculation yourself.