Jeyanthi Thangiah

Designing Virtual Team Members

The intern model: agents as long-running specialists with a defined scope, one VM each, improving themselves over time.

Jeyanthi Thangiah2 min read

The intern model treats an agent as a specialist with a well-defined scope, rather than as a general assistant.

This article expands on the agent-design theme in my AI Engineer World’s Fair 2026 overview. It focuses on OpenRouter’s conference presentation about treating agents as virtual team members, with defined responsibilities and their own working environments.

The short version

  • The intern model: AI agents as specialists with well-defined scope — long-running, living in Slack, self-improving
  • A concrete build recipe: git-as-config, one VM per intern, credentials kept outside model-visible context, skill marketplaces
  • Fleet counts reported by OpenRouter in its presentation (63 FTEs, 73 interns, 368 unique skills) and the model-selection lesson learned

Shashank Goyal (OpenRouter) gave the most operationally concrete talk of the conference: "Letting the Interns Loose," about building the machine that builds OpenRouter. The framing: AI agents as interns — specialists with well-defined scope, because "constraints make agents reliable" under a narrower set of operating conditions; constraints alone do not establish determinism.

"Constraints make agents reliable." — Shashank Goyal, OpenRouter

Three properties define an intern:

  • It is long-running
  • It lives in Slack ("talk where you are" — this fixed the low employee-adoption problem)
  • It improves itself

The build recipe:

Recipe stepHow it works
Every intern is a git repoConfig as code, skills and workflows in files, one source of truth, version-controlled audit logs. Identity, prompts, rules, and schedules all live in files; the harness (Claude Agent SDK reading a .claude directory of skills, plus webhooks) reads from code.
One VM per intern (diagram above)Durable local state for memory management, hard isolation, runnable from anywhere (a single EC2 instance serving many users, or GitHub Actions), and simple to reason about.
Keep credentials outside model-visible contextIn the reported design, the host control plane supplies credentials to tool execution. Host access is different from model access: tool output, logs, files, or an unrestricted shell can still expose credentials unless those paths are constrained and tested.
Learning and reuseInterns learn skills, which flow into an intern skill marketplace across the fleet.

Ori, the intern manager, applies biological versioning: shared DNA (a base template), fork your own, evolve independently. The presentation reported these counts, not measured productivity gains: 63 FTEs, 73 interns, 368 unique skills across the fleet. Key lesson learned: model preference is personal to each intern — start with the smartest model, then use evals to downgrade to cheaper ones where quality holds. Cookbooks: agent harness TUI, long-horizon agents.

Try this next

  • Pilot one narrowly scoped intern as a git repo — identity, prompts, rules, and schedules in files — running on a single dedicated EC2 instance or AgentCore Runtime session for hard isolation
  • Keep credentials in the control plane or narrowly scoped tool execution; test that model-visible files, shell access, outputs, and logs cannot disclose them
  • Start each intern on the smartest available Bedrock model, then use evals to downgrade to cheaper models where quality holds

Continue in

AI Engineer World's Fair 2026

Agentic AI

AI Engineer World's Fair 2026

Four days, 300 speakers, 39 tracks. Five themes carried across all of it: harness engineering and software factories, agent loops, evals and observability, the unsettled argument about agent memory, and the shift from burning tokens to proving value.

Read post
AI Innovation

Build Loops That Compound

Verification is what separates a loop that compounds from one that merely repeats — Sonar's case for verified agent loops, a control-theory recipe for loop design, and how to bound loops that spin out.

Read post
AI Innovation

Evals Done Right

Improving an agent is a data-mining problem: collect traces, curate them, and let agents read the traces for you. Plus what it takes to survive the eval rollercoaster.

Read post