Jeyanthi Thangiah

Sandboxes: More Autonomy, Smaller Blast Radius

Approval loops exist because we don't trust what an agent might do to the host. Containerise it and gateway its tool access, and the blast radius becomes bounded — isolation is what purchases autonomy.

Jeyanthi Thangiah4 min read

How much autonomy an agent can safely be given is a function of its blast radius. Sandboxing is how that radius gets small enough to stop asking permission for every action.

This article explores the execution boundaries behind the agent systems discussed in my AI Engineer World’s Fair 2026 overview. Drawing on the conference’s workshops and talks, it looks at how isolated environments and controlled tool access can give agents more room to work while limiting the damage a mistake can cause.

The short version

  • Why containers are the default agent sandbox: bounded, auditable blast radius is the engineering path from approval loops to autonomy
  • The managed-agents progression — Messages API → Agent SDK → managed agents — and which harness pieces the platform absorbs at each step
  • The light/heavy split: the speaker described roughly 70% lightweight work in the workload under discussion; this is not a general population estimate

Containers as the Default Sandbox

Docker's workshop (ai-engineer-workshop) charted the path "from approval loops to autonomous agents": the container is the natural agent sandbox, and Docker's agent stack — compose-for-agents for declaring multi-agent systems, plus the MCP toolkit, cagent, and gateway for curated, containerized tool access (multi-agent build guide) — makes agent environments as declarative and reproducible as application deployment.

The framing in the workshop title is the argument: approval loops exist because we don't trust what the agent might do to the host. Containerize the agent and gateway its MCP tool access, and the blast radius becomes bounded and auditable — the more trustworthy the isolation, the fewer human approval gates you need. That is the actual engineering path from approval loops to autonomy, and it lands on the same conclusion as OpenRouter's one-VM-per-intern and Antigravity's per-sub-agent sandboxing: isolation is what purchases autonomy.

Isolation is what purchases autonomy. — the shared conclusion of Docker's workshop, OpenRouter's one-VM-per-intern, and Antigravity's per-sub-agent sandboxing

Managed Agents: The Platform Runs the Sandbox

Anthropic's session traced three generations of agentic surface, each absorbing more harness into the platform:

GenerationWhat you getWhat you still build
Claude 3 / Messages APITokens in, tokens outThe agentic loop, tool calling, context management, and all production infra (sessions, observability, credentials, hosting) from scratch
Claude Agent SDKA built-in agentic loop (the Claude Code harness) with filesystem accessSession management, credentials, hosting, and execution isolation are still yours
Claude managed agentsCall the "brain" and get the loop, the sandbox, and all production infra run by Anthropic—

The engineering principles behind managed agents are broadly instructive. Build for the model capabilities of tomorrow — "context anxiety" (Sonnet 4.5 rushing as the window fills; Opus doesn't need the fix) shows that when the model moves and the harness doesn't, it degrades the agent.

When the model moves and the harness doesn't, it degrades the agent

  • "Context anxiety" is the worked example: Sonnet 4.5 rushes as the window fills and needs a harness fix; Opus doesn't
  • A harness tuned to yesterday's model quirks becomes a drag on tomorrow's model — build for the capabilities of tomorrow

Decouple the brain from the hands: separate the agent loop from the tool-execution environment so you can spin up sessions on demand, retry a new session if one dies, and persist context in durable memory. The core primitives are Agent × Environment × Sessions (a durable agent instance); reliability is modeled into the session state machine; and the context window is explicitly not the session — everything logs to a persistent session log. The production recipe demonstrated (an SRE agent): define the agent → create an environment → provide evidence (logs, files) → run sessions of agent × environment × evidence.

Light Sandboxes: Most Agent Code Doesn't Need a VM

Samuel Colvin's contribution is the light/heavy split: he described an approximately 70/30 split for the workload under discussion, not a universal distribution of agent execution. Evaluate your own tasks and isolation requirements before choosing a substrate. Pydantic Monty (announcement, Simon Willison's review, HN, community experience report) is a purpose-built lightweight sandbox for exactly that class of task, paired with the pydantic-ai-harness and hardened via a public hack-Monty challenge. Design rules regardless of substrate: a curated list of tools with a small surface area, scoped to specific tasks, not broad tasks.

Try this next

  • Default agent execution into isolated containers or AgentCore Runtime sessions with gatewayed MCP tool access (e.g. AgentCore Gateway), so the blast radius is bounded and approval gates can be relaxed deliberately
  • Apply the light/heavy split: measure the lightweight share of your workload, then select an execution environment that satisfies the required isolation for each task
  • Persist every session to a durable log independent of the context window, and model retries/failures in an explicit session state machine

AI Engineer World's Fair 2026

Agentic AI

AI Engineer World's Fair 2026

Four days, 300 speakers, 39 tracks. Five themes carried across all of it: harness engineering and software factories, agent loops, evals and observability, the unsettled argument about agent memory, and the shift from burning tokens to proving value.

Read post
AI Innovation

Build Loops That Compound

Verification is what separates a loop that compounds from one that merely repeats — Sonar's case for verified agent loops, a control-theory recipe for loop design, and how to bound loops that spin out.

Read post
AI Innovation

Evals Done Right

Improving an agent is a data-mining problem: collect traces, curate them, and let agents read the traces for you. Plus what it takes to survive the eval rollercoaster.

Read post