Sandboxes: More Autonomy, Smaller Blast Radius
Approval loops exist because we don't trust what an agent might do to the host. Containerise it and gateway its tool access, and the blast radius becomes bounded — isolation is what purchases autonomy.

How much autonomy an agent can safely be given is a function of its blast radius. Sandboxing is how that radius gets small enough to stop asking permission for every action.
This article explores the execution boundaries behind the agent systems discussed in my AI Engineer World’s Fair 2026 overview. Drawing on the conference’s workshops and talks, it looks at how isolated environments and controlled tool access can give agents more room to work while limiting the damage a mistake can cause.
The short version
- Why containers are the default agent sandbox: bounded, auditable blast radius is the engineering path from approval loops to autonomy
- The managed-agents progression — Messages API → Agent SDK → managed agents — and which harness pieces the platform absorbs at each step
- The light/heavy split: the speaker described roughly 70% lightweight work in the workload under discussion; this is not a general population estimate
Containers as the Default Sandbox
Docker's workshop (ai-engineer-workshop) charted the path "from approval loops to autonomous agents": the container is the natural agent sandbox, and Docker's agent stack — compose-for-agents for declaring multi-agent systems, plus the MCP toolkit, cagent, and gateway for curated, containerized tool access (multi-agent build guide) — makes agent environments as declarative and reproducible as application deployment.
The framing in the workshop title is the argument: approval loops exist because we don't trust what the agent might do to the host. Containerize the agent and gateway its MCP tool access, and the blast radius becomes bounded and auditable — the more trustworthy the isolation, the fewer human approval gates you need. That is the actual engineering path from approval loops to autonomy, and it lands on the same conclusion as OpenRouter's one-VM-per-intern and Antigravity's per-sub-agent sandboxing: isolation is what purchases autonomy.
Isolation is what purchases autonomy. — the shared conclusion of Docker's workshop, OpenRouter's one-VM-per-intern, and Antigravity's per-sub-agent sandboxing
Managed Agents: The Platform Runs the Sandbox
Anthropic's session traced three generations of agentic surface, each absorbing more harness into the platform:
| Generation | What you get | What you still build |
|---|---|---|
| Claude 3 / Messages API | Tokens in, tokens out | The agentic loop, tool calling, context management, and all production infra (sessions, observability, credentials, hosting) from scratch |
| Claude Agent SDK | A built-in agentic loop (the Claude Code harness) with filesystem access | Session management, credentials, hosting, and execution isolation are still yours |
| Claude managed agents | Call the "brain" and get the loop, the sandbox, and all production infra run by Anthropic | — |
The engineering principles behind managed agents are broadly instructive. Build for the model capabilities of tomorrow — "context anxiety" (Sonnet 4.5 rushing as the window fills; Opus doesn't need the fix) shows that when the model moves and the harness doesn't, it degrades the agent.
When the model moves and the harness doesn't, it degrades the agent
- "Context anxiety" is the worked example: Sonnet 4.5 rushes as the window fills and needs a harness fix; Opus doesn't
- A harness tuned to yesterday's model quirks becomes a drag on tomorrow's model — build for the capabilities of tomorrow
Decouple the brain from the hands: separate the agent loop from the tool-execution environment so you can spin up sessions on demand, retry a new session if one dies, and persist context in durable memory. The core primitives are Agent × Environment × Sessions (a durable agent instance); reliability is modeled into the session state machine; and the context window is explicitly not the session — everything logs to a persistent session log. The production recipe demonstrated (an SRE agent): define the agent → create an environment → provide evidence (logs, files) → run sessions of agent × environment × evidence.
Light Sandboxes: Most Agent Code Doesn't Need a VM
Samuel Colvin's contribution is the light/heavy split: he described an approximately 70/30 split for the workload under discussion, not a universal distribution of agent execution. Evaluate your own tasks and isolation requirements before choosing a substrate. Pydantic Monty (announcement, Simon Willison's review, HN, community experience report) is a purpose-built lightweight sandbox for exactly that class of task, paired with the pydantic-ai-harness and hardened via a public hack-Monty challenge. Design rules regardless of substrate: a curated list of tools with a small surface area, scoped to specific tasks, not broad tasks.
Try this next
- Default agent execution into isolated containers or AgentCore Runtime sessions with gatewayed MCP tool access (e.g. AgentCore Gateway), so the blast radius is bounded and approval gates can be relaxed deliberately
- Apply the light/heavy split: measure the lightweight share of your workload, then select an execution environment that satisfies the required isolation for each task
- Persist every session to a durable log independent of the context window, and model retries/failures in an explicit session state machine
Continue in



