Contextual RAG
How adding document context to chunks can improve retrieval, what Anthropic actually measured, and how to evaluate the tradeoffs.

A passage can be perfectly clear in its document and almost meaningless in a search result. “This policy applies from next quarter” tells us little when we cannot see which policy, whose quarter, or which version of the document it came from.
Contextual retrieval addresses that problem before search: add a short explanation of the chunk’s document context, then index the enriched text.
What the benchmark actually says
Anthropic’s September 2024 evaluation reported these top-20 retrieval failure rates, using its selected embedding configuration across several datasets:
| Configuration | Failure rate |
|---|---|
| Baseline embeddings | 5.7% |
| Contextual Embeddings | 3.7% |
| Contextual Embeddings + Contextual BM25 | 2.9% |
| Contextual Embeddings + Contextual BM25 + reranking | 1.9% |
The last change is 3.8 percentage points, or about 66.7% relative reduction. As an arithmetic illustration, 57 failures per 1,000 comparable retrieval opportunities would become 19: 38 fewer. It does not mean 67 fewer wrong answers per 100 questions. The metric is retrieval failure, based on recall@20, rather than answer accuracy or patient outcomes. These are company-reported results, not measurements of a system I deployed. Anthropic’s methodology and results.
The 2.9% row includes contextualization in both indexes. Calling it ordinary hybrid search would credit the wrong intervention. Similarly, the 1.9% result includes reranking; contextual embeddings alone did not produce it.
A concrete example
Consider this fictional maintenance document:
- Title: Cooling System B — Maintenance Policy
- Effective date: January 2025
- Section: Filter inspection
- Chunk: “Inspect it every month. Replace it if the seal is damaged.”
A useful prefix would be: “This passage concerns the filter in Cooling System B under the maintenance policy effective January 2025.” Store that generated prefix separately from the original quotation. A reader should be able to distinguish source text from an interpretation added during indexing.
The contextualizer needs the surrounding document or an appropriate source window. A prompt with only the isolated chunk cannot reliably recover a missing antecedent. Ask it to use only supplied source information, preserve dates and entity names, and leave unresolved references unresolved. Then inspect a sample for invented context. Enrichment can introduce errors as well as resolve ambiguity.
A proposed pipeline
- Parse documents while retaining titles, section boundaries, versions, and permissions.
- Split them into chunks and preserve links to the original passages.
- Generate a short contextual prefix using the relevant document context.
- Index the combined text for vector search and lexical search.
- Retrieve authorized candidates, combine the rankings, and optionally rerank.
- Give the generator the selected passages, distinguishing original text from generated context.
- Return citations that open the original version and passage.
This is a reference design. It is not evidence that a particular combination of cloud services has passed deployment, latency, or security testing.
What it costs
Context generation happens during indexing, but that does not establish zero effect on query latency. Larger retrieved passages affect generation; reranking adds a runtime operation. Updates can also require regenerating context and embeddings.
A useful cost model separates input tokens, output tokens, cached reads and writes, embeddings, storage, retrieval, and reranking. For example, assuming 50,000 chunks, 2,000 billed input tokens per chunk and an illustrative input rate of $0.15 per million tokens:
50,000 × 2,000 ÷ 1,000,000 × $0.15 = $15
That is input processing alone, not $0.15 total and not a complete indexing bill. At 100 output tokens per chunk and an illustrative $0.60 per million output tokens, output adds $3. These rates are arithmetic assumptions, not a current model price quote. Whole-document prompts may have much larger inputs; caching changes the accounting. Price a representative run with actual token usage before extrapolating.
Decide from the failure, not the headline
Use a held-out set of questions to compare ordinary hybrid retrieval with contextual retrieval, holding other components steady. Inspect retrieval recall, answer support, citation correctness, refusal when evidence is absent, and performance on outdated or conflicting documents. Report sample sizes, labeling methods, model versions, cost, and latency distributions.
This approach is most worth investigating when missing document context is a recurring cause of failure. It may offer little when well-structured passages already name their entities and dates. Fixing parsing or chunk boundaries may solve the problem more simply.
Legal and clinical applications still need domain-specific validation, appropriate human review, and a way to stop or escalate unsupported answers. A retrieval benchmark does not define an acceptable clinical error rate.
The useful promise is specific: make a retrieved passage easier to identify and interpret. Whether that improves the final decision is something the evaluation still has to establish.
Continue in


