AI Engineer World's Fair 2026
Four days, 300 speakers, 39 tracks. Five themes carried across all of it: harness engineering and software factories, agent loops, evals and observability, the unsettled argument about agent memory, and the shift from burning tokens to proving value.
- speakers
- 300
- attendees
- 6,000+
- tracks
- 39
- days
- 4
What this post covers
- Agent harness engineering + software factoriesThe headline theme of the whole conference
- Loops (+ traces, long-running/multi-agent)Sensor → controller → actuator, borrowed from control theory
- Evals & observabilityThe verifiers are king — and observability is how you find out at scale
- Agent memory & context engineering (+ skills)Nobody's won this argument yet
- Model economics / tokenomics / local LLMs2025 was "burn tokens fast." 2026 is "prove the value."
- Also at the conferenceThe other 34 tracks, in one slide
- Sources · everything linkedRead the originals, watch the talks

Moscone West, 28 June to 2 July 2026. Three hundred speakers, more than 6,000 attendees, four days of tracks — Day 2 was "Software Factories," Day 3 "Autoresearch," Day 4 "Harness Engineering." These are the five themes that carried across all of it.
"Every time we made it easier to write software, we ended up writing exponentially more of it. The future belongs to engineers who make agent work legible, verifiable, and worth shipping."
— Addy Osmani, on a Day 3 slide
Agent harness engineering + software factories
The headline theme of the whole conference
What is a Harness
Harness = the engineering scaffolding you build around a swappable model: memory, retrieval, tool surface, sandboxes, and the agent loop itself. The model is the frozen reasoning core; everything else is yours to design. Day 4's keynote track was literally named "Harness Engineering" — this was the word of the conference.
Agent = model + harness. The harness is context, control, action, and persistence wired around the model chip — with failure feeding back into new rules via hooks and a ratchet.
Anthropic
"Decouple the brain from the hands" — the model reasons, the harness executes and sandboxes.
AWS
The harness as a "production cage" — the scaffolding that makes an agent safe to run unattended.
LangChain
Formalized it as an equation: Agent = Model + Harness. Same model, wildly different agent.
Lay of the land
Every vendor's product maps onto the same decomposition: model → storage → retrieval → memory → semantic layer → agent loop → context engineering (Oracle's canonical breakdown). The industry term for "an assembly line that stamps out reliable agents from shared harness infrastructure" is agent factory — the frame for both use cases below.


Use case 1 — Notion's Software Factory
Notion's "Model Agnostic Playbook" — model-agnostic by design, auto-router handles 75% of traffic on the moderate (non-frontier) path. The screenshots below are a real Notion product doc, not a diagram — the concrete "agents run the inner loop, humans run the outer loop" example.



The takeaway: the human (a PM named Geoffrey, in this doc) decided what to build and when to ship. The agents did the building and the reviewing. That split — not the tooling — is the pattern worth remembering.
Use case 2 — OpenRouter's interns
Where Notion's factory faces users, OpenRouter's faces the company itself: 73 agent "interns," each a self-contained git repo — config-as-code, skills and workflows as one source of truth — working alongside 63 human FTEs.
Y Combinator's Garry Tan makes the same point from the employer side: "workforce made of markdown" — an AI-native org treats agent instructions as versioned files, same discipline as code.
The tool surface is part of the harness — MCP Apps
A harness decides not just which tools an agent can reach but what a tool is allowed to look like. MCP Apps (SEP-1865, from Ido Salomon and Liad Yosef — spec authors, not framework users) standardizes interactive UI inside the MCP host, so a tool can return a real interface instead of a wall of text. The framing: agents are becoming the new browsers, and the web wasn't built for them.
Bot traffic has overtaken human traffic
And agents largely ignore
llms.txt-style files — they reach for docs and homepages instead. Properly linked agent-ready files see 4× the usage.The agentic web will never be fully headless
Some moments — comparing options, picking a seat — still need a human. MCP Apps is the last mile between agent and user, not a replacement for it.
The counterpoint — why software factories fail
Dex Horthy (HumanLayer, who coined "context engineering") gave the most useful talk against the headline theme: harness engineering is necessary but not sufficient. A factory automates triage → plan → build → test → review → ship → monitor — and the failure mode is model slop: low-quality output that compounds in large codebases. His CTO's companion talk put it bluntly: "bad code compounded, and agents created problems that agents couldn't solve — until we had to throw it all away."
Model slop
Models aren't trained for maintainability at scale. Automate-everything pipelines produce code that compounds into unmaintainable systems.
Benchmark inadequacy
Evals measure "did tests pass," never "is this maintainable." The gap the industry can't currently see, let alone fix.
No human steering
Scheduled loops without correction drift from quality over time. Unbounded "loop until solved" agents lose coherence and spin out — repeating a broken approach without noticing.
The bar for high-stakes actions: for writing to private data or communicating on a user's behalf, 90% accuracy is not acceptable. As the 12-factor-agents README puts it — "can you imagine a web app that crashed on 10% of page loads?" One mistake on an irreversible action is permanent.
Agents own the inner loop. Humans own the outer loop.
The pattern that recurs in every topic today
Read more
Where I go deeper on this: Building an Agent Factory, Designing Virtual Team Members, Sandboxes — Room Without Risk, Agent Frameworks Compared.
Loops (+ traces, long-running/multi-agent)
Sensor → controller → actuator, borrowed from control theory
What a loop is
A loop = sensor (measures the gap) → controller (decides the next small change) → actuator (applies it). Addy Osmani's version of the same primitive: loop = goal + cadence + isolated work + verification + state — five parts, and the recursive goal iterates until done, but the outer verdict is still an engineering decision.


Set point
The desired end state you're steering toward.
Sensor
Measures the gap between current and desired state.
Controller
Decides the next small, low-risk change.
Actuator
The agent that applies the change and opens a PR.
- Kyle Mistele, HumanLayer — 4:43 control theory (17:57 total)
- Lance Martin, Anthropic — 07:23 build/verify loop primitive
- Anthropic — Getting Started with Loops
- Anthropic — Startup Builds: Getting Started with Loops (webinar)
- Anthropic cookbook — Verify with an outcome grader
- Anthropic — Startup Builds webinar recording (on-demand)
Anthropic frames the same shift as three decisions you make before a loop runs — and getting them wrong is now a more common failure than writing a bad prompt.

Types of loops
LangSmith's Sydney Reimer breaks the single idea above into four loops of increasing scope — each one wraps the previous:
Four loops
Loop 1 — the core agent loop. Model receives context, calls tools, tools return observations, agent continues until complete.





The anti-pattern, named for a reason: a "Ralph Wiggum loop" tries things, fails, and keeps going without noticing. Loops without verification don't compound — they just repeat, badly. A scheduled loop without human steering drifts; verification (Topic 3) is what makes a loop compound instead of merely repeat.


Continuous improvement from traces
LangChain's framing: "improving agents is a data mining problem." Build v1 → run it → collect traces → mine them → hill-climb. Traces run millions of tokens long — too long for humans to read, so the fix is to send an agent to read the traces, with rubrics for what to look for (bad interactions, cost, compaction behavior).
Open models are viable as trace judges. The question isn't "use the best model," it's "what's the cheapest model that matches Opus's trace-judging capability for this task?"
Long-running / multi-agent — the substrate that makes loops possible
OpenRouter's interns live in Slack with durable state — loops need somewhere to persist between runs. Google Antigravity's lead agent dynamically builds its own sub-agent graph at runtime (not a static DAG the developer wired up) — loops that restructure themselves per task.
You can't have a compounding loop without persistent state, and you can't scale a loop without spreading it across agents — which is why loops and long-running/multi-agent are one topic here.
Anthropic's own answer to "how do you actually build this": three concrete harness primitives that keep a long session honest across restarts.
Default-FAIL contract
Every criterion starts
false; a hook denies marking it passing until the agent has actually opened the evidence.Fresh-context evaluator
A separate agent with no write access grades the work from a context window that never saw the build — the builder doesn't grade its own work.
Agent-maintained handoff
The agent writes its own progress notes and commits to git, so the next session — which has no memory of the last one — picks up cleanly.
What engineers own vs. what's autonomous
The inner-loop/outer-loop split from Topic 1 has a concrete boundary: evidence. The agent's inner loop — investigate, implement, test, report — is capability. The engineer's outer loop — decide, verify, approve, own — is agency. Evidence (diffs, tests, logs, why) is what crosses from one side to the other.


Osmani's agency ladder makes "high agency" concrete instead of a hiring buzzword. Five rungs, low to high:
Flag
Notice a problem and leave it for the system.
Execute
Do the work yourself.
Diagnose / propose / recommend
Understand the problem and put forward a fix, without necessarily owning execution.
Resolve
Take it all the way to done and own the outcome.
Discernment (rare, top rung)
Find a problem and decide whether it's even worth investing in. When agents make more paths possible, agency isn't chasing every one — it's deciding which deserve your ownership.
Read more
Where I go deeper on this: Build Loops That Compound, Building a Multi-Agent Orchestration System on AWS.
Evals & observability
The verifiers are king — and observability is how you find out at scale
Evals — the glance
"In the land of AI agents, the verifiers are king" — Tariq Shaukat, CEO, Sonar. Loop engineering without verification is just automation.




SonderMind's mental-health AI coach is the sharpest worked example of harness-as-eval-infrastructure from the whole conference. The design principle is modularity: an input-middleware pipeline — guardrails, memory, personalization, history — sits in front of the model, with a three-tier routing layer calibrated jointly with clinicians. Guardrails are a first-class harness component, not a model setting: API-hosted models' built-in guardrails actively get in the way for sensitive mental-health scenarios, so the team turns them off and builds their own, calibrated and testable.
"The hardest safety signals are indirect, coded, and emerge across turns."
— SonderMind
Five generalizable harness-design rules fall out of this and recur across every eval-mature system at the conference:
Make every layer independently testable
SonderMind's modularity is precisely what makes their eval gating possible — calibrate the input guardrail without touching routing or memory.
Keep the model swappable
Oracle's frozen reasoning core; OpenRouter's per-intern model choice with eval-driven downgrades.
Curate the tool surface
A curated tool set beats a buffet of generic MCPs — small surface area for specific tasks.
Log to a durable session record, independent of the context window
The context window is a view, not the source of truth.
Put verification inside the loop, not just at the end
In-loop verification (Sonar's SonarVertex) catches issues as the agent works, when correction is cheapest.
Observability — where we land
Evals tell you if the agent is good; observability/tracing is how you find out at scale and feed it back. Long-horizon agents produce millions of tokens of trace data — no human reads that. You need a mining/judging pipeline over the traces before you can even ask "is this getting better?"
LangSmith
Deepest LangChain/LangGraph integration. The LangSmith Engine mines traces into training signal — the reference architecture for closing the loop: trace mining → distillation/fine-tuning; evals + env-gen.
Braintrust
Eval-first. Logging is secondary to scoring and comparison — the inverse emphasis from LangSmith: optimizes the eval-iteration loop itself.
Langfuse
What we actually run. OSS, MIT-licensed, native OTel, self-hosted. Borrow the patterns from the platforms above, not the platform itself.
Close the loop: this observability layer is why Topic 2's "continuous improvement from traces" pattern (the one we just covered) was possible at all.
Where I go deeper on this: Evals Done Right.
Agent memory & context engineering (+ skills)
Nobody's won this argument yet
The industry is split on how agents should remember and retrieve. Both camps below have real production examples from major labs at the same conference — this is not a recommendation, it's the honest state of the debate.
Two camps
- Vector similarity finds "similar." It does not find related.
- Multi-hop reasoning ("Drug A → side effect B → contraindicated with Drug C") needs graph traversal, not embeddings
- Neo4j's workshop demo: entity extraction + relationship-building alongside the vector index
- Atlan's production example: internal skills chained as a context graph — "competitive intel skill" feeds "category positioning skill"
- Oracle's take goes furthest: a single governed, unified memory core in-database, not a bolt-on vector store next to everything else
- Honest limitation: typedef built a full graph context layer — and found the agent mostly grepped the filesystem anyway
- Agents defined by files, not embeddings — OpenRouter's 368 shared skills, each a markdown file
- YC's "workforce made of markdown" (Garry Tan) — one SKILL.md per agent employee
- Anthropic showed Claude Sonnet 3.5 writing raw files to a memory directory — no vector DB, "that's really it"
- Honest limitation: early memories are shallow ("caterpie and weedle are both caterpillar Pokémon") — this camp still struggles with depth
Camp A — structured memory, further out
Oracle takes this camp furthest: a single governed, unified memory core built into the database itself, not a vector store bolted onto everything else — one substrate for episodic, semantic, and procedural memory instead of a dozen disconnected services. On the research side, Graph RAG keeps pushing the same multi-hop argument from Camp A's bullets above.
The learning flywheel pattern
Oracle's promotion pipeline made concrete: agent completes a task → the workflow is captured → recurring workflows (3+ times in 30 days) get promoted to a SHA-versioned skill → future sessions retrieve it by intent. The agent teaches itself its own most common tasks — and it needs a structured store to do it.
The most fully-realized version of this camp came from typedef, who call it the Data Context Layer: "a semantic ontology of every definition, pipeline, dashboard, and system across your data platform, compiled from your code, always current, always traceable."

The most honest finding in this whole topic — and it comes from inside Camp A. Their design principles were exactly what this camp argues for: deterministic first, LLM only where it earns its keep; edges carry provenance and confidence; computed once on commit; "traversal beats prompt stuffing." Their conclusion after building it: the agent didn't traverse the graph nearly as much as they'd hoped — it grepped the filesystem anyway. Worth sitting with before we invest in graph infrastructure ourselves.
Camp B — filesystem memory, live
"You basically give Claude memory tools... the ability to write to a file system, that is basically a memory directory. That's really it."
— Lance Martin, Anthropic
Modeled on human memory: short-term/experiential (hippocampus) vs. long-term consolidated (cortex, written during "dreaming"). Compare to the community consensus we found in a widely-shared Reddit thread: naive/vanilla RAG is dead; hybrid isn't. Nobody's saying structure is unnecessary — the fight is over how much structure, and where it lives.
What actually shipped in personal-assistant memory
Shlok Khemani surveyed every shipped personal-assistant memory system through mid-2026 — ChatGPT, Claude, both iterating hard. The counterintuitive finding: none of them use RAG for core memory. RAG retrieves content; personal memory needs continuous synthesis — tracking preferences that change, coherent across dozens of conversations. What every system converged on instead is the running profile: a periodically-synthesized document, not a database of facts.


raw/ holds whatever you dumped in unedited, alongside a wikis/ tree organized by entity type (people/, concepts/, sources/). Right: after an enrich-note agent runs on a cadence — tags added, source and enrichment timestamp recorded, and entities extracted so it can wire nvidia-news to gpu-prices as related. The visualizations/ folder is the maintenance layer: a connection graph and a burn-down chart of stale or under-linked notes.The Karpathy LLM Wiki pattern
A personal knowledge base with a graph view of connected notes and a note burn-down — surfacing stale or under-linked notes for review before the base silently degrades.
The knowledge base pattern: raw notes → living wikis
Raw capture is fast and unstructured; an enrichment agent adds tags, source, and entity extraction on a schedule, turning noise into a base the agent can actually retrieve from. Ben Holmes (Warp) — the talk this diagram is from.
The practitioner-experience angle
Anthropic's Fable field guide talk (Thariq Shihipar) is worth watching in full — a rare look at what changed in the harness as the underlying model got smarter across Opus 4 → 4.5 → 4.8, not just the memory/context-anxiety angle below.


implementation-notes.md file. If you hit an edge case that forces you to deviate from the plan, pick the conservative option, log it under 'Deviations,' and keep going." — the same agent-maintained-handoff pattern from Topic 2, in Fable's own words.

AskUserQuestion; Opus 4.5 could interview you; Opus 4.8 could build the interview itself. The harness didn't change — the model's use of it did.Worth a full watch, not just the clips above: the whole progression of the Fable harness across model generations is one of the clearest "the harness gets simpler as the model gets smarter" stories from the conference.
This one stays an open question. Anyone doing agent memory work has to pick a lane, and the industry has genuinely not settled which one is right.
Read more
Where I go deeper on this: Context Engineering.
Model economics / tokenomics / local LLMs
2025 was "burn tokens fast." 2026 is "prove the value."
Where costs have been headed
Era 1 (2025) — "tokenmaxxing": burn tokens in parallel to compress a week of work into a day. Real productivity gain, real cost. Era 2 (2026) — "intelligence per dollar": the metric shifted from tokens-burned to value-produced-per-dollar.

The routing decision that matters most
Notion's playbook, open weights
Notion's 5-point Model Agnostic Playbook: build for multi-model by default, evaluate on value not tokens, switch every 2–3 weeks, give frontier labs feedback, forgo discounts for optionality.



Local AI
The NVIDIA/DGX Spark panel argued local AI has crossed from "interesting" to "useful" — sovereignty, regulated industries, and now genuinely cost-competitive.
The routing rule to take away: start on Haiku, escalate to Sonnet only when Haiku fails eval, use Opus only for subtasks that demonstrably require it — with the justification documented in code. Defaulting to Sonnet "just to be safe" costs 4× more; defaulting to Opus costs 15× more.
Read more
Where I go deeper on this: The Token Scarcity Era, How Model Routing Works, Making Applications Model-Agnostic by Design.
Also at the conference
The other 34 tracks, in one slide
The conference ran 39 tracks over four days. The five topics above are the ones worth going deepest on. Here is what else was significant, so none of it looks like it was not there.
Sandboxing & agent security
Named the hardest unsolved problem of H2 2026: the "lethal trifecta" — untrusted content + private data + external comms — with no mature solution yet. Pydantic's useful split: ~70% of agent code execution is millisecond-scale and needs a lightweight interpreter (Monty), not a VM.
Microsoft Foundry & the "IQ" layer
Microsoft's keynote pitch: retrieval as infrastructure, not app code. Foundry IQ plus Work/Fabric/Web IQ as one entry point into an organization's ambient data — agentic retrieval, semantic ranker, MCP built in. The closest thing to an enterprise-platform answer to the memory debate.
Local AI & inference
A whole Day 4 track. NVIDIA's panel argued local has crossed from interesting to useful — sovereignty, regulated industries, cost, resilience — on DGX Spark-class hardware with multi-node clustering. Inference was named the dominant cost line in Amplify's State of AI Engineering.
Robotics & world models
Its own Day 3 track (alongside Computer Use) — the physical-embodiment end of agents. A different problem space from most agent engineering, but it is where the long-horizon planning research is being pushed hardest.
Post-training & mid-training
A full Day 3 track on the layer below the harness: what you do to a model between pretraining and serving it. Relevant context for why "the harness gets simpler as the model gets smarter" keeps happening.
Domain tracks we skipped
AI in Healthcare, AI in Finance (JPMorgan's learned execution graphs), Agentic Commerce, Voice & Realtime, Vision & OCR, Generative Media, Design Engineering, Data Quality. Healthcare is the one worth a follow-up given who we are.
One more worth naming — because it's the path we're already on. Vercel's internal AI data scientist answers ~1,200 unique queries a day and went through six-plus rewrites to get there. The sequence: one mega-prompt → separate agents with dedicated system prompts → embedding Claude Code → a filesystem agent → adding skills → a harness framework. That last rewrite became Eve, now open source (runtime, durable workflow, sandbox, observability built in). The migration path is the takeaway, not the framework.
Sources · everything linked
Read the originals, watch the talks
Watch these talks
- Notion's Token Town — Sarah Sachs (23:55)
- Loop Engineering from First Principles — HumanLayer (17:57)
- Claude for Long-Horizon Tasks — Lance Martin (25:18)
- In the Land of AI Agents — Sonar (18:53)
- The Great Loops Debate (1:00:16)
- Every company should have a Brain — Garry Tan (~20:36)
- Field Guide to Fable — Thariq Shihipar (19:28)
- WTF Is the Context Layer — Prukalpa Sankar (20:53)
- Local AI State of the Union (44:29)
- Evals-Driven Development — SonderMind (21:17)
Read these
- The 2026 AI Engineering Report — Amplify Partners
- Sonar — Loop engineering without verification
- Simon Willison — The lethal trifecta
- LangChain — Improving Agents is a Data Mining Problem
- Geoff Huntley — Ralph Wiggum as a software engineer
- Anthropic — Scaling Managed Agents
- Tomasz Tunguz — Tokenmaxxing
- Langfuse
- AI21 — Structured RAG
Continue in


