Jeyanthi Thangiah

Building an Agent Factory

SonderMind's input-middleware pipeline with three-tier routing, reported in a mental-health deployment, and the harness rules that generalise from it.

Jeyanthi Thangiah3 min read

A safety-first harness puts middleware in front of the model — guardrails, memory, personalization, history — rather than treating safety as a setting on the model.

The short version

  • A safety-first harness architecture (SonderMind's input-middleware pipeline with three-tier routing) reported in a mental-health deployment
  • Five generalizable harness-design rules: independently testable layers, swappable models, curated tool surface, durable session logs, in-loop verification
  • The load-bearing factory split — agents run the inner loop, humans own the outer loop — and the failure modes of factories that skip it

Designing the Harness: Safety-First Middleware

SonderMind’s presentation describes a mental-health AI coach and its harness design. A harness is the surrounding software that manages model inputs, tools, state, and execution. This is a reported case study, not an independent clinical safety evaluation. The design principle is modularity: an input middleware pipeline — input guardrails, memory, personalization, history — sits in front of the model, with a three-tier routing layer calibrated jointly with clinicians deciding how each message is handled. Guardrails are a first-class harness component, not a model setting: the presenters described provider filters that interfered with their particular workflow and the custom safeguards they calibrated. That account is not a general instruction to disable provider safeguards; any deployment needs evidence that its full safety system handles the intended clinical scenarios. The hardest safety signals are "indirect, coded, and emerge across turns" — a reason to evaluate conversation-level signals as well as per-message filters; neither guarantees detection.

The hardest safety signals are "indirect, coded, and emerge across turns." — SonderMind

Generalizable harness-design rules from across the sessions:

  • Make every layer independently testable — SonderMind's modularity is precisely what makes their eval gating possible; you can calibrate the input guardrail without touching routing or memory
  • Keep the model swappable — Oracle's frozen reasoning core; OpenRouter's per-intern model choice with eval-driven downgrades
  • Curate the tool surface — typedef's "curated tool set versus a buffet of generic MCPs"; Pydantic's small surface area for specific tasks
  • Log everything to a durable session record independent of the context window — Anthropic's managed agents persist a session log so the context window is a view, not the source of truth
  • Put verification inside the loop, not just at the end — Sonar's in-loop verification (SonarVertex) catches issues as the agent works, when correction is cheapest

Agents Run the Inner Loop, Humans Own the Outer Loop

The inner-loop / outer-loop factory split

A software factory is an agentic pipeline automating the full development loop: triage → plan → build → test → review → ship → monitor. The load-bearing design decision is the split shown above, articulated identically by steipete and HumanLayer: agents own the inner loop; humans own the outer loop. In the Openclaw pattern, an agent files an issue, a manager agent spawns workers that understand the product's goals and intent, agents work autonomously, humans review and approve, and the loop continues — agent continues the inner loop, humans control the outer loop.

The human side of the outer loop has its own discipline — decide, verify, approve, own — and an agency ladder for how much initiative the agent (or engineer) takes:

The agency ladder

Dex Horthy's talk supplied the failure analysis for factories that skip this.

Three ways factories fail without human steering

  • Model slop — automate-everything pipelines produce unmaintainable code at scale
  • Benchmark inadequacy — evals measure "tests pass," not maintainability
  • No human steering — scheduled loops without correction drift

His three-part fix:

  • Model-assisted preplanning and alignment — program design upfront, vertical slices, use the model to plan not just execute
  • Quality evals that measure maintainability — SWE-Marathon, DeepSWE, Frontier Code, fastlane, abundant.ai
  • Model + harness co-ownership

Industry case studies: StrongDM's software factory, factory.ai, Faros.ai.

Try this next

  • Structure your Bedrock/AgentCore harness as independently testable middleware layers (guardrails, memory, routing) so each can be calibrated and eval-gated without touching the others
  • Keep the model swappable behind Bedrock's model abstraction and log every session to a durable store independent of the context window
  • Define your outer loop explicitly: which pipeline stages agents own end-to-end, and where humans decide, verify, approve, and own

Keep reading

Agentic AI

AI Engineer World's Fair 2026

Four days, 300 speakers, 39 tracks. Five themes carried across all of it: harness engineering and software factories, agent loops, evals and observability, the unsettled argument about agent memory, and the shift from burning tokens to proving value.

Read post