Evals Done Right
Improving an agent is a data-mining problem: collect traces, curate them, and let agents read the traces for you. Plus what it takes to survive the eval rollercoaster.

Improving an agent is mostly a data-mining problem: run a v1, collect traces, curate them, and run experiments on what the traces show.
This is a deeper look at the evaluation and observability theme in my AI Engineer World’s Fair 2026 overview. It brings together the conference sessions on execution traces, evaluation design, and expert feedback to explore how teams can tell whether an agent is improving.
The short version
- The trace-mining loop: build a v1 agent, collect traces, curate them, run experiments — with agents reading the traces for you
- Google's sequencing advice for surviving the eval rollercoaster: start early, start small, test the negatives
- A high-stakes case study (SonderMind) of evals gating releases, and the expert-in-the-loop flywheel that scales domain judgment
Improving Agents Is a Data-Mining Problem
"Improving agents is a data mining problem." — Vivek Trivedy, LangChain
Vivek Trivedy (LangChain) frames the loop: build a v1 agent and start running it → collect a ton of traces → curate the trace data (this is the data-mining step) → run experiments on your data. Because long-horizon traces run to millions of tokens, you centralize tracing data and send an agent to read the traces with explicit instructions — find bad interactions, check compaction behavior after transactions, compute cost — and rubrics to codify what to look for. Open models are cost-efficient trace judges: find the minimal intelligence level that matches Opus's trace-judging capability for your specific task.
The LangSmith Engine productizes the loop: trace mining feeding distillation and fine-tuning; evals and environment generation ("define agent behavior by showing the evals we ran on it" — the agent hill-climbs the evals); and product analytics for high-trust domains. The deeper frame is model–harness–task fit — fit(model, harness, task), where Agent = Model + Harness and Task = Data — and the cycle is harness engineering → fine-tuning → harness engineering again, collecting feedback as fast as possible. Continual-learning frontier: the agent acts, updates information about itself, and "sleep-time compute and dreaming" turn accumulated traces into training data, harness updates, and memory. See also: Improving Deep Agents with Harness Engineering and the referenced paper.
Surviving the Eval Rollercoaster: Start Small, Test Negatives
Google's production-evals team (Daniel Bump, Preetika Bhateja, Chris Souza) delivered pragmatic sequencing advice built around the eval rollercoaster: early on, intuition-based validation is fine (it's not scalable, but it's fast); during rapid prototyping, heavyweight evals can hinder progress because single prompt tweaks produce large gains; the mistake is scaling evals too early, before the eval itself is calibrated. Start early, start small — you don't need a massive golden set on day 1 — and test the negatives: check that the model didn't do something bad, not just that it did the right thing.
Eval rollercoaster failure modes
- Scaling evals too early — before the eval itself is calibrated — is the core mistake; during rapid prototyping, heavyweight evals can hinder progress because single prompt tweaks produce large gains
- Watch for overfitting: agents that optimize for a narrow dataset fail to generalize
- Don't trust isolated runs — rely on patterns across runs (how often does a task pass?)
For scaling:
- Clear rubric templates with trusted examples
- Explanations alongside verdicts — multi-output ("Is it accurate? Is it brand safe?") beats bare pass/fail, and explanations make the agent better
- Monitor LLM-judge alignment against human ground truth
- Spot-check reasoning and intent via agent tracing
- Curate the golden set deliberately
Watch for overfitting — agents that optimize for a narrow dataset fail to generalize. Hill-climbing is rewarding when you iterate on the eval and the agent together. For launch readiness: understand regressions, rely on patterns across runs (how often does a task pass?) rather than isolated runs, and invest in online evals. Their definition of a good eval system:
| Property | What it means |
|---|---|
| Representative | Covers what the product must be great at |
| Evolving | Grows with new user patterns |
| Expert-audited golden set | Includes edge cases |
| Comprehensive templates/rubrics | With clear examples |
| Clear gate-keeping rule | Everyone knows what blocks a release |
Case Study: Guardrails Under Real Safety Stakes
SonderMind's session was the conference's best case study in evals-driven development under genuine safety stakes (a mental-health AI coach). The central problem: "the hardest signals are indirect, coded, and emerge across turns" — crisis indicators that no per-message filter catches. Their answer: guardrail thresholds and the three-tier routing are calibrated with clinicians, clinicians annotate evals directly in LangSmith (adding "expected observation" labels), annotations become tight labeled scenarios (clinical category + trigger), and evals gate releases — catching both false positives and false negatives.
Overcalibration is a real failure mode
- "Panic systems don't help" — a guardrail that over-triggers degrades the product it protects
- Yet when forced to choose, SonderMind treats over-calibration as "a compassionate choice": the false-positive/false-negative tradeoff is an explicit ethical decision, not a tuning detail
Two hard-won judgments stand out. First, overcalibration is a real failure mode — "panic systems don't help" — yet when forced to choose, they treat over-calibration as "a compassionate choice": the false-positive/false-negative tradeoff is an explicit ethical decision, not a tuning detail. Second, the presenters described conflicts between provider filters and legitimate conversations in their specific workflow. Their custom-safeguard approach is a case-specific engineering choice, not a general requirement to disable provider safeguards. It needs clinical oversight and evaluation of the combined safety system, including missed crises and inappropriate escalation. SonderMind open-sourced their guardrail calibration datasets (GitHub — datasets only; the guardrail prompts are not included).
The Expert Is the Center of the Loop

The flywheel above generalizes SonderMind's deepest structural insight: the human expert's judgment sits in CI — the human is the center node of the learning loop. Clinicians don't review outputs after the fact; their annotated judgments are converted into eval cases that run on every release. This is the same principle as HumanLayer's outer-loop steering and Google's human-ground-truth alignment monitoring, applied at maximum stakes: the domain expert is encoded into the pipeline, so their judgment scales with the agent rather than bottlenecking it. The general recipe: identify whose judgment defines "good" in your domain, give them a low-friction annotation surface, compile annotations into gating evals, and route disagreements between automated judges and the expert back into calibration.
Try this next
- Centralize agent traces (e.g. via your observability stack on Bedrock/AgentCore) and send an agent to mine them with explicit instructions and rubrics; find the cheapest Bedrock model that matches your top model's trace-judging quality
- Start a small golden set now — including negative tests that check the agent didn't do something bad — and grow it with expert-audited edge cases before scaling eval infrastructure
- Identify whose judgment defines "good" in each domain, give them a low-friction annotation surface, and compile their annotations into release-gating evals
Continue in



