Jeyanthi Thangiah

AI Engineer World's Fair 2026

Four days, 300 speakers, 39 tracks. Five themes carried across all of it: harness engineering and software factories, agent loops, evals and observability, the unsettled argument about agent memory, and the shift from burning tokens to proving value.

Jeyanthi Thangiah23 min read
speakers
300
attendees
6,000+
tracks
39
days
4

Moscone West, 28 June to 2 July 2026. Three hundred speakers, more than 6,000 attendees, four days of tracks — Day 2 was "Software Factories," Day 3 "Autoresearch," Day 4 "Harness Engineering." These are the five themes that carried across all of it.

"Every time we made it easier to write software, we ended up writing exponentially more of it. The future belongs to engineers who make agent work legible, verifiable, and worth shipping."

— Addy Osmani, on a Day 3 slide

Agent harness engineering + software factories

The headline theme of the whole conference

What is a Harness

Harness = the engineering scaffolding you build around a swappable model: memory, retrieval, tool surface, sandboxes, and the agent loop itself. The model is the frozen reasoning core; everything else is yours to design. Day 4's keynote track was literally named "Harness Engineering" — this was the word of the conference.

Agent = model + harness. The harness is context, control, action, and persistence wired around the model chip — with failure feeding back into new rules via hooks and a ratchet.

Anthropic

"Decouple the brain from the hands" — the model reasons, the harness executes and sandboxes.

Managed Agents engineering post →

AWS

The harness as a "production cage" — the scaffolding that makes an agent safe to run unattended.

Bedrock AgentCore →

LangChain

Formalized it as an equation: Agent = Model + Harness. Same model, wildly different agent.

Improving Deep Agents with Harness Engineering →

Lay of the land

Every vendor's product maps onto the same decomposition: model → storage → retrieval → memory → semantic layer → agent loop → context engineering (Oracle's canonical breakdown). The industry term for "an assembly line that stamps out reliable agents from shared harness infrastructure" is agent factory — the frame for both use cases below.

Agentic software factory — the win is moving human judgment to the highest leverage checkpoint
Product intent, incidents, and user feedback feed the agent inner loop; evidence crosses the boundary into human verdict. "The win is not removing people from the loop — it's moving human judgment to the highest leverage checkpoint."
A closer look at software factories — taskboard, triage, specs, build, review, ship, monitor
The factory floor in detail: an agent triages, drafts specs, implements, and monitors — with explicit human-review gates on specs and implementation before anything ships.

Use case 1 — Notion's Software Factory

Notion's "Model Agnostic Playbook" — model-agnostic by design, auto-router handles 75% of traffic on the moderate (non-frontier) path. The screenshots below are a real Notion product doc, not a diagram — the concrete "agents run the inner loop, humans run the outer loop" example.

Decagon Duet agent summarizing user research on deleted blocks
1. Research surfaces the problem — the Decagon Duet agent mines support tickets: 71% of complaints were about verifying deleted content.
Claude Agent writes a PR
2. Claude Agent writes the fix — a real PR, tagged for review by Codex Agent.
Codex Agent reviews and catches two edge cases
3. Codex Agent reviews it — catches 2 edge cases, pushes fixes before merge.

The takeaway: the human (a PM named Geoffrey, in this doc) decided what to build and when to ship. The agents did the building and the reviewing. That split — not the tooling — is the pattern worth remembering.

Use case 2 — OpenRouter's interns

Where Notion's factory faces users, OpenRouter's faces the company itself: 73 agent "interns," each a self-contained git repo — config-as-code, skills and workflows as one source of truth — working alongside 63 human FTEs.

73agent "interns," each a git repo with its own skills & workflows
368unique skills shared across the fleet — a skill marketplace, not a shared prompt
63human FTEs working alongside the fleet, not replaced by it

Y Combinator's Garry Tan makes the same point from the employer side: "workforce made of markdown" — an AI-native org treats agent instructions as versioned files, same discipline as code.

The tool surface is part of the harness — MCP Apps

A harness decides not just which tools an agent can reach but what a tool is allowed to look like. MCP Apps (SEP-1865, from Ido Salomon and Liad Yosef — spec authors, not framework users) standardizes interactive UI inside the MCP host, so a tool can return a real interface instead of a wall of text. The framing: agents are becoming the new browsers, and the web wasn't built for them.

  • Bot traffic has overtaken human traffic

    And agents largely ignore llms.txt-style files — they reach for docs and homepages instead. Properly linked agent-ready files see 4× the usage.

  • The agentic web will never be fully headless

    Some moments — comparing options, picking a seat — still need a human. MCP Apps is the last mile between agent and user, not a replacement for it.

The counterpoint — why software factories fail

Dex Horthy (HumanLayer, who coined "context engineering") gave the most useful talk against the headline theme: harness engineering is necessary but not sufficient. A factory automates triage → plan → build → test → review → ship → monitor — and the failure mode is model slop: low-quality output that compounds in large codebases. His CTO's companion talk put it bluntly: "bad code compounded, and agents created problems that agents couldn't solve — until we had to throw it all away."

  1. Model slop

    Models aren't trained for maintainability at scale. Automate-everything pipelines produce code that compounds into unmaintainable systems.

  2. Benchmark inadequacy

    Evals measure "did tests pass," never "is this maintainable." The gap the industry can't currently see, let alone fix.

  3. No human steering

    Scheduled loops without correction drift from quality over time. Unbounded "loop until solved" agents lose coherence and spin out — repeating a broken approach without noticing.

The bar for high-stakes actions: for writing to private data or communicating on a user's behalf, 90% accuracy is not acceptable. As the 12-factor-agents README puts it — "can you imagine a web app that crashed on 10% of page loads?" One mistake on an irreversible action is permanent.

Agents own the inner loop. Humans own the outer loop.

The pattern that recurs in every topic today

Loops (+ traces, long-running/multi-agent)

Sensor → controller → actuator, borrowed from control theory

What a loop is

A loop = sensor (measures the gap) → controller (decides the next small change) → actuator (applies it). Addy Osmani's version of the same primitive: loop = goal + cadence + isolated work + verification + state — five parts, and the recursive goal iterates until done, but the outer verdict is still an engineering decision.

Applications of Agentic Control Loops — eradicate bad patterns, adopt frameworks incrementally, maintain a fork, mirror to another language, keep integrations up to date, ensure compliance, resolve system anomalies
Kyle Mistele — what this primitive is actually for: eradicating bad patterns, incremental framework adoption, maintaining a fork against upstream, mirroring a project into another language, keeping integrations current, compliance, and resolving system anomalies.
Agentic Control Loops diagram — Desired End State, Controller, Actuator (Agent), System (Codebase), Sensor with eslint rules, react-doctor, agent+skill, ast-grep, rg, packwerk
Kyle Mistele, HumanLayer — the exact loop primitive every "agentic loop" talk assumed you already knew. The sensor isn't abstract: eslint rules, ast-grep, ripgrep, or another agent+skill, watching for violations in the codebase.

Set point

The desired end state you're steering toward.

Sensor

Measures the gap between current and desired state.

Controller

Decides the next small, low-risk change.

Actuator

The agent that applies the change and opens a PR.

Anthropic frames the same shift as three decisions you make before a loop runs — and getting them wrong is now a more common failure than writing a bad prompt.

Anthropic — Loop engineering: three decisions, when it starts, what stops it if it stalls, who decides it's finished
Anthropic — when it starts (prompt, goal, timer, event), what stops it if it stalls (turns, time, or you cancelling it), who decides it's finished (a test, a schema, a script — with nothing to check, only Claude's judgement ends the loop).

Types of loops

LangSmith's Sydney Reimer breaks the single idea above into four loops of increasing scope — each one wraps the previous:

Four loops

Loop 1 — the core agent loop. Model receives context, calls tools, tools return observations, agent continues until complete.

The verification loop — Loop 2
Loop 2 — the verification loop. A grader scores the result against criteria; unmet criteria feed back into the agent loop.
The event-driven loop — Loop 3
Loop 3 — the event-driven loop. Schedule triggers, webhooks, and Slack messages kick off the verified agent loop automatically — this is what makes agents always-on.
The hill climbing loop — Loop 4
Loop 4 — the hill climbing loop. Production traces become the signal for what's working, what's failing, and what should change in the harness itself.
Human oversight fits in every loop — before sensitive tool calls, as the grader, on workflow results, and reviewing harness changes
Putting it all together — the four loops summary
The advantage isn't just in the agent you build — it's in the loops you build around it.

The anti-pattern, named for a reason: a "Ralph Wiggum loop" tries things, fails, and keeps going without noticing. Loops without verification don't compound — they just repeat, badly. A scheduled loop without human steering drifts; verification (Topic 3) is what makes a loop compound instead of merely repeat.

The Ralph technique — while :; do cat PROMPT.md | claude-code; done
Geoff Huntley's Ralph technique (Jul 2025) — put the agent in an infinite loop, one task per iteration. "Deterministically bad in an undeterministic world."
Claude Code implementation of the Ralph loop with a Stop hook
A Stop hook + iteration guard is the same loop, without the shell wrapper — this is how you bound a Ralph loop safely inside Claude Code.

Continuous improvement from traces

LangChain's framing: "improving agents is a data mining problem." Build v1 → run it → collect traces → mine them → hill-climb. Traces run millions of tokens long — too long for humans to read, so the fix is to send an agent to read the traces, with rubrics for what to look for (bad interactions, cost, compaction behavior).

Open models are viable as trace judges. The question isn't "use the best model," it's "what's the cheapest model that matches Opus's trace-judging capability for this task?"

Long-running / multi-agent — the substrate that makes loops possible

OpenRouter's interns live in Slack with durable state — loops need somewhere to persist between runs. Google Antigravity's lead agent dynamically builds its own sub-agent graph at runtime (not a static DAG the developer wired up) — loops that restructure themselves per task.

You can't have a compounding loop without persistent state, and you can't scale a loop without spreading it across agents — which is why loops and long-running/multi-agent are one topic here.

Anthropic's own answer to "how do you actually build this": three concrete harness primitives that keep a long session honest across restarts.

  1. Default-FAIL contract

    Every criterion starts false; a hook denies marking it passing until the agent has actually opened the evidence.

  2. Fresh-context evaluator

    A separate agent with no write access grades the work from a context window that never saw the build — the builder doesn't grade its own work.

  3. Agent-maintained handoff

    The agent writes its own progress notes and commits to git, so the next session — which has no memory of the last one — picks up cleanly.

What engineers own vs. what's autonomous

The inner-loop/outer-loop split from Topic 1 has a concrete boundary: evidence. The agent's inner loop — investigate, implement, test, report — is capability. The engineer's outer loop — decide, verify, approve, own — is agency. Evidence (diffs, tests, logs, why) is what crosses from one side to the other.

The loop boundary is evidence — agent inner loop vs engineer outer loop
Agents run capability loops. Engineers own agency loops. The boundary between them is evidence — not trust, not vibes.
Agents run the inner loop. Engineers own the outer loop.
Restated plainly: agents run the inner loop, engineers own the outer loop.

Osmani's agency ladder makes "high agency" concrete instead of a hiring buzzword. Five rungs, low to high:

  1. Flag

    Notice a problem and leave it for the system.

  2. Execute

    Do the work yourself.

  3. Diagnose / propose / recommend

    Understand the problem and put forward a fix, without necessarily owning execution.

  4. Resolve

    Take it all the way to done and own the outcome.

  5. Discernment (rare, top rung)

    Find a problem and decide whether it's even worth investing in. When agents make more paths possible, agency isn't chasing every one — it's deciding which deserve your ownership.

Evals & observability

The verifiers are king — and observability is how you find out at scale

Evals — the glance

"In the land of AI agents, the verifiers are king" — Tariq Shaukat, CEO, Sonar. Loop engineering without verification is just automation.

44%fewer AI-driven production outages from zero-trust, multi-layer verification
92%reduction in issues — one bank's test of Sonar's guide-verify-solve loop
3distinct failure types Arize maps separately: deterministic, hallucination, trajectory
👩‍⚕️SonderMind puts clinician judgment directly in CI — best worked example of the pattern
Sonar ACDC verification loops diagram — Master the critical verification loops to bring the Agent Centric Development Cycle to life
Sonar's ACDC loop — three nested loops (agentic, CI verification, code maintenance), each built on "guide, verify, solve." Neglect verification and the same loop that compounds gains compounds a downward spiral just as fast.
Arize different evals for different failures
Arize — different failure types need different eval types: code evaluators, LLM-as-judge, agent-as-judge.
Arize AX — financial_agent trace tree with a failed QA evaluation flagged incorrect on a risky day-trading response
Arize's actual debugging view: a financial-agent trace tree, span by span (authenticate → risk assessment → compliance screening → response generation), with an online eval flagged incorrect — the agent suggested risky day-trading with an emergency fund, and the eval explains exactly why that's wrong. This is the "eval as data on the trace itself" pattern — an AI-judge layer running continuously against production traces, not a one-off test suite.
SonderMind Agent Harness — actual conference slide, input/output guardrails, memory, personalization, history, Sonder + Coach LLM
SonderMind's actual conference slide. Input/output guardrail middleware in front of Sonder and the Coach LLM — "every architectural decision was made with safety as the primary objective." The hardest safety signals are indirect and emerge across turns — exactly what a single-message filter can't catch.

SonderMind's mental-health AI coach is the sharpest worked example of harness-as-eval-infrastructure from the whole conference. The design principle is modularity: an input-middleware pipeline — guardrails, memory, personalization, history — sits in front of the model, with a three-tier routing layer calibrated jointly with clinicians. Guardrails are a first-class harness component, not a model setting: API-hosted models' built-in guardrails actively get in the way for sensitive mental-health scenarios, so the team turns them off and builds their own, calibrated and testable.

"The hardest safety signals are indirect, coded, and emerge across turns."

— SonderMind

Five generalizable harness-design rules fall out of this and recur across every eval-mature system at the conference:

  1. Make every layer independently testable

    SonderMind's modularity is precisely what makes their eval gating possible — calibrate the input guardrail without touching routing or memory.

  2. Keep the model swappable

    Oracle's frozen reasoning core; OpenRouter's per-intern model choice with eval-driven downgrades.

  3. Curate the tool surface

    A curated tool set beats a buffet of generic MCPs — small surface area for specific tasks.

  4. Log to a durable session record, independent of the context window

    The context window is a view, not the source of truth.

  5. Put verification inside the loop, not just at the end

    In-loop verification (Sonar's SonarVertex) catches issues as the agent works, when correction is cheapest.

Observability — where we land

Evals tell you if the agent is good; observability/tracing is how you find out at scale and feed it back. Long-horizon agents produce millions of tokens of trace data — no human reads that. You need a mining/judging pipeline over the traces before you can even ask "is this getting better?"

LangSmith

Deepest LangChain/LangGraph integration. The LangSmith Engine mines traces into training signal — the reference architecture for closing the loop: trace mining → distillation/fine-tuning; evals + env-gen.

LangSmith Engine docs →

Braintrust

Eval-first. Logging is secondary to scoring and comparison — the inverse emphasis from LangSmith: optimizes the eval-iteration loop itself.

braintrust.dev →

Langfuse

What we actually run. OSS, MIT-licensed, native OTel, self-hosted. Borrow the patterns from the platforms above, not the platform itself.

langfuse.com →

Close the loop: this observability layer is why Topic 2's "continuous improvement from traces" pattern (the one we just covered) was possible at all.

Where I go deeper on this: Evals Done Right.

Agent memory & context engineering (+ skills)

Nobody's won this argument yet

The industry is split on how agents should remember and retrieve. Both camps below have real production examples from major labs at the same conference — this is not a recommendation, it's the honest state of the debate.

Two camps

  • Vector similarity finds "similar." It does not find related.
  • Multi-hop reasoning ("Drug A → side effect B → contraindicated with Drug C") needs graph traversal, not embeddings
  • Neo4j's workshop demo: entity extraction + relationship-building alongside the vector index
  • Atlan's production example: internal skills chained as a context graph — "competitive intel skill" feeds "category positioning skill"
  • Oracle's take goes furthest: a single governed, unified memory core in-database, not a bolt-on vector store next to everything else
  • Honest limitation: typedef built a full graph context layer — and found the agent mostly grepped the filesystem anyway
  • Agents defined by files, not embeddings — OpenRouter's 368 shared skills, each a markdown file
  • YC's "workforce made of markdown" (Garry Tan) — one SKILL.md per agent employee
  • Anthropic showed Claude Sonnet 3.5 writing raw files to a memory directory — no vector DB, "that's really it"
  • Honest limitation: early memories are shallow ("caterpie and weedle are both caterpillar Pokémon") — this camp still struggles with depth

Camp A — structured memory, further out

Oracle takes this camp furthest: a single governed, unified memory core built into the database itself, not a vector store bolted onto everything else — one substrate for episodic, semantic, and procedural memory instead of a dozen disconnected services. On the research side, Graph RAG keeps pushing the same multi-hop argument from Camp A's bullets above.

  • The learning flywheel pattern

    Oracle's promotion pipeline made concrete: agent completes a task → the workflow is captured → recurring workflows (3+ times in 30 days) get promoted to a SHA-versioned skill → future sessions retrieve it by intent. The agent teaches itself its own most common tasks — and it needs a structured store to do it.

    Read more →

The most fully-realized version of this camp came from typedef, who call it the Data Context Layer: "a semantic ontology of every definition, pipeline, dashboard, and system across your data platform, compiled from your code, always current, always traceable."

typedef's Data Context Layer — a data engineer asks what breaks if we enable multi-currency in Salesforce, agents traverse a graph spanning online, streaming, offline and BI systems, returning an impact report
The worked example that sells the idea: a data engineer asks "what breaks if we enable multi-currency in Salesforce?" — agents traverse the context layer across online, streaming, offline, and BI systems and return "24 models + 3 reports affected," with an action log showing every step. Pre-analyzed understanding of your entire data system: how data flows, what it means, and what breaks when things change.

The most honest finding in this whole topic — and it comes from inside Camp A. Their design principles were exactly what this camp argues for: deterministic first, LLM only where it earns its keep; edges carry provenance and confidence; computed once on commit; "traversal beats prompt stuffing." Their conclusion after building it: the agent didn't traverse the graph nearly as much as they'd hoped — it grepped the filesystem anyway. Worth sitting with before we invest in graph infrastructure ourselves.

Camp B — filesystem memory, live

"You basically give Claude memory tools... the ability to write to a file system, that is basically a memory directory. That's really it."

— Lance Martin, Anthropic

Modeled on human memory: short-term/experiential (hippocampus) vs. long-term consolidated (cortex, written during "dreaming"). Compare to the community consensus we found in a widely-shared Reddit thread: naive/vanilla RAG is dead; hybrid isn't. Nobody's saying structure is unnecessary — the fight is over how much structure, and where it lives.

What actually shipped in personal-assistant memory

Shlok Khemani surveyed every shipped personal-assistant memory system through mid-2026 — ChatGPT, Claude, both iterating hard. The counterintuitive finding: none of them use RAG for core memory. RAG retrieves content; personal memory needs continuous synthesis — tracking preferences that change, coherent across dozens of conversations. What every system converged on instead is the running profile: a periodically-synthesized document, not a database of facts.

Memory is a function of compute — consolidation cost vs profile-size cost
Two cost dials: cost to maintain ∝ frequency × intelligence of the consolidating model; cost to serve ∝ profile size injected into every new chat.
Warp knowledge base folder structure — the same raw/ folder before and after an enrich-note agent runs, showing added tags and a related link between two notes
Before and after the same folder. Left: raw/ holds whatever you dumped in unedited, alongside a wikis/ tree organized by entity type (people/, concepts/, sources/). Right: after an enrich-note agent runs on a cadence — tags added, source and enrichment timestamp recorded, and entities extracted so it can wire nvidia-news to gpu-prices as related. The visualizations/ folder is the maintenance layer: a connection graph and a burn-down chart of stale or under-linked notes.
  • The Karpathy LLM Wiki pattern

    A personal knowledge base with a graph view of connected notes and a note burn-down — surfacing stale or under-linked notes for review before the base silently degrades.

    Read more →

  • The knowledge base pattern: raw notes → living wikis

    Raw capture is fast and unstructured; an enrichment agent adds tags, source, and entity extraction on a schedule, turning noise into a base the agent can actually retrieve from. Ben Holmes (Warp) — the talk this diagram is from.

    Read more →

The practitioner-experience angle

Anthropic's Fable field guide talk (Thariq Shihipar) is worth watching in full — a rare look at what changed in the harness as the underlying model got smarter across Opus 4 → 4.5 → 4.8, not just the memory/context-anxiety angle below.

The most important part of working with Fable is staying in the loop
"The most important part of working with Fable is staying in the loop."
Fable implementation notes — log deviations from the plan and keep going
"Keep an implementation-notes.md file. If you hit an edge case that forces you to deviate from the plan, pick the conservative option, log it under 'Deviations,' and keep going." — the same agent-maintained-handoff pattern from Topic 2, in Fable's own words.
System Prompt Design evolution — small system prompt few tools lots of examples, to large system prompt lots of examples many tools, to smaller system prompt with tool search and no examples
System prompt design converged the opposite way from where most teams start: smaller system prompt, tool search instead of a big tool list, no few-shot examples.
AskUserQuestion capability progression across Opus 4, 4.5, and 4.8 — could call it, could interview you, could build the interview
The same tool, three model generations: Opus 4 could call AskUserQuestion; Opus 4.5 could interview you; Opus 4.8 could build the interview itself. The harness didn't change — the model's use of it did.

Worth a full watch, not just the clips above: the whole progression of the Fable harness across model generations is one of the clearest "the harness gets simpler as the model gets smarter" stories from the conference.

This one stays an open question. Anyone doing agent memory work has to pick a lane, and the industry has genuinely not settled which one is right.

Where I go deeper on this: Context Engineering.

Model economics / tokenomics / local LLMs

2025 was "burn tokens fast." 2026 is "prove the value."

Where costs have been headed

Era 1 (2025) — "tokenmaxxing": burn tokens in parallel to compress a week of work into a day. Real productivity gain, real cost. Era 2 (2026) — "intelligence per dollar": the metric shifted from tokens-burned to value-produced-per-dollar.

$300MSalesforce's annual Anthropic token spend — while freezing eng headcount
4 mohow long Uber ran unconstrained AI spend before capping it
~$1.3Mannual savings from a 10% routing shift at Salesforce's spend level
AI value maturity: Thought Partner, Assistant, Teammates, The System
Value climbs with automation level — Thought Partner → Assistant → Teammates → The System — but so does the bill.

The routing decision that matters most

1×Haiku — routing, extraction, formatting
4×Sonnet — multi-step reasoning, code generation
15×Opus — novel reasoning, long-horizon planning
75%of Notion's autonomous-agent traffic still runs on the non-frontier path

Notion's playbook, open weights

Notion's 5-point Model Agnostic Playbook: build for multi-model by default, evaluate on value not tokens, switch every 2–3 weeks, give frontier labs feedback, forgo discounts for optionality.

Notion's Model Agnostic Playbook, 5 points
Notion's 5-point playbook — model choice as an ongoing decision, not a one-time bet.
Open-weight models are now strong enough to handle these workloads
Open weights caught up — a cost lever and negotiating leverage against frontier pricing.
Baseten model capabilities chart, closed vs open
Baseten — closed vs. open model capability curves on real SWE tasks, over time.

Local AI

The NVIDIA/DGX Spark panel argued local AI has crossed from "interesting" to "useful" — sovereignty, regulated industries, and now genuinely cost-competitive.

The routing rule to take away: start on Haiku, escalate to Sonnet only when Haiku fails eval, use Opus only for subtasks that demonstrably require it — with the justification documented in code. Defaulting to Sonnet "just to be safe" costs 4× more; defaulting to Opus costs 15× more.

Also at the conference

The other 34 tracks, in one slide

The conference ran 39 tracks over four days. The five topics above are the ones worth going deepest on. Here is what else was significant, so none of it looks like it was not there.

Sandboxing & agent security

Named the hardest unsolved problem of H2 2026: the "lethal trifecta" — untrusted content + private data + external comms — with no mature solution yet. Pydantic's useful split: ~70% of agent code execution is millisecond-scale and needs a lightweight interpreter (Monty), not a VM.

Simon Willison — the lethal trifecta →

Microsoft Foundry & the "IQ" layer

Microsoft's keynote pitch: retrieval as infrastructure, not app code. Foundry IQ plus Work/Fabric/Web IQ as one entry point into an organization's ambient data — agentic retrieval, semantic ranker, MCP built in. The closest thing to an enterprise-platform answer to the memory debate.

What is Foundry IQ →

Local AI & inference

A whole Day 4 track. NVIDIA's panel argued local has crossed from interesting to useful — sovereignty, regulated industries, cost, resilience — on DGX Spark-class hardware with multi-node clustering. Inference was named the dominant cost line in Amplify's State of AI Engineering.

DGX Spark →

Robotics & world models

Its own Day 3 track (alongside Computer Use) — the physical-embodiment end of agents. A different problem space from most agent engineering, but it is where the long-horizon planning research is being pushed hardest.

Post-training & mid-training

A full Day 3 track on the layer below the harness: what you do to a model between pretraining and serving it. Relevant context for why "the harness gets simpler as the model gets smarter" keeps happening.

Domain tracks we skipped

AI in Healthcare, AI in Finance (JPMorgan's learned execution graphs), Agentic Commerce, Voice & Realtime, Vision & OCR, Generative Media, Design Engineering, Data Quality. Healthcare is the one worth a follow-up given who we are.

One more worth naming — because it's the path we're already on. Vercel's internal AI data scientist answers ~1,200 unique queries a day and went through six-plus rewrites to get there. The sequence: one mega-prompt → separate agents with dedicated system prompts → embedding Claude Code → a filesystem agent → adding skills → a harness framework. That last rewrite became Eve, now open source (runtime, durable workflow, sandbox, observability built in). The migration path is the takeaway, not the framework.

Sources · everything linked

Read the originals, watch the talks

Continue in

Keep reading