Building an Agent Factory
SonderMind's input-middleware pipeline with three-tier routing, reported in a mental-health deployment, and the harness rules that generalise from it.

A safety-first harness puts middleware in front of the model — guardrails, memory, personalization, history — rather than treating safety as a setting on the model.
The short version
- A safety-first harness architecture (SonderMind's input-middleware pipeline with three-tier routing) reported in a mental-health deployment
- Five generalizable harness-design rules: independently testable layers, swappable models, curated tool surface, durable session logs, in-loop verification
- The load-bearing factory split — agents run the inner loop, humans own the outer loop — and the failure modes of factories that skip it
Designing the Harness: Safety-First Middleware
SonderMind’s presentation describes a mental-health AI coach and its harness design. A harness is the surrounding software that manages model inputs, tools, state, and execution. This is a reported case study, not an independent clinical safety evaluation. The design principle is modularity: an input middleware pipeline — input guardrails, memory, personalization, history — sits in front of the model, with a three-tier routing layer calibrated jointly with clinicians deciding how each message is handled. Guardrails are a first-class harness component, not a model setting: the presenters described provider filters that interfered with their particular workflow and the custom safeguards they calibrated. That account is not a general instruction to disable provider safeguards; any deployment needs evidence that its full safety system handles the intended clinical scenarios. The hardest safety signals are "indirect, coded, and emerge across turns" — a reason to evaluate conversation-level signals as well as per-message filters; neither guarantees detection.
The hardest safety signals are "indirect, coded, and emerge across turns." — SonderMind
Generalizable harness-design rules from across the sessions:
- Make every layer independently testable — SonderMind's modularity is precisely what makes their eval gating possible; you can calibrate the input guardrail without touching routing or memory
- Keep the model swappable — Oracle's frozen reasoning core; OpenRouter's per-intern model choice with eval-driven downgrades
- Curate the tool surface — typedef's "curated tool set versus a buffet of generic MCPs"; Pydantic's small surface area for specific tasks
- Log everything to a durable session record independent of the context window — Anthropic's managed agents persist a session log so the context window is a view, not the source of truth
- Put verification inside the loop, not just at the end — Sonar's in-loop verification (SonarVertex) catches issues as the agent works, when correction is cheapest
Agents Run the Inner Loop, Humans Own the Outer Loop

A software factory is an agentic pipeline automating the full development loop: triage → plan → build → test → review → ship → monitor. The load-bearing design decision is the split shown above, articulated identically by steipete and HumanLayer: agents own the inner loop; humans own the outer loop. In the Openclaw pattern, an agent files an issue, a manager agent spawns workers that understand the product's goals and intent, agents work autonomously, humans review and approve, and the loop continues — agent continues the inner loop, humans control the outer loop.
The human side of the outer loop has its own discipline — decide, verify, approve, own — and an agency ladder for how much initiative the agent (or engineer) takes:

Dex Horthy's talk supplied the failure analysis for factories that skip this.
Three ways factories fail without human steering
- Model slop — automate-everything pipelines produce unmaintainable code at scale
- Benchmark inadequacy — evals measure "tests pass," not maintainability
- No human steering — scheduled loops without correction drift
His three-part fix:
- Model-assisted preplanning and alignment — program design upfront, vertical slices, use the model to plan not just execute
- Quality evals that measure maintainability — SWE-Marathon, DeepSWE, Frontier Code, fastlane, abundant.ai
- Model + harness co-ownership
Industry case studies: StrongDM's software factory, factory.ai, Faros.ai.
Try this next
- Structure your Bedrock/AgentCore harness as independently testable middleware layers (guardrails, memory, routing) so each can be calibrated and eval-gated without touching the others
- Keep the model swappable behind Bedrock's model abstraction and log every session to a durable store independent of the context window
- Define your outer loop explicitly: which pipeline stages agents own end-to-end, and where humans decide, verify, approve, and own
Continue in


