Jeyanthi Thangiah

Graph RAG

How graph traversal can supply relationship evidence, what the AWS/Lettria benchmark establishes, and how to bound a proposed implementation.

Jeyanthi Thangiah4 min read

Some questions depend on relationships: which components depend on a retired service, which contracts refer to an amended clause, or which cases cite a decision that was later overturned? A graph makes selected connections explicit so retrieval can follow them.

Vector retrieval can also support multi-hop answers, especially with iterative search. A graph is a different representation and retrieval strategy, not proof that every alternative is structurally incapable of reasoning.

What the cited result establishes

The December 2024 AWS/Lettria article reports 80% correct answers for its graph approach versus 50.83% for its comparison RAG system. That is 29.17 percentage points, approximately 57% relative improvement, or 1.57 times the baseline—not 3.2 times.

The evaluation was conducted in-house across several domains, including finance, healthcare research, aeronautics, and environmental regulation. It is vendor-reported evidence from that setup. It does not substantiate the law-firm outcome, attorney satisfaction scores, or ROI figures previously presented here. It also does not establish clinical safety. AWS/Lettria report and evaluation description.

A fictional example with explicit relationships

Suppose a small dependency graph records:

  • Service A depends on Library B.
  • Library B depends on Package C.
  • Package C is scheduled for retirement.

A traversal can find Service A through two dependency edges. The graph has made the chain directly queryable. But if the Library B → Package C edge is missing, the answer may still be incomplete. If an extractor confuses two packages with the same name, an apparently clear path may be wrong.

Every extracted relationship should retain its source passage, version, and relevant dates. Keep uncertain extractions distinguishable from reviewed facts. A graph path is inspectable evidence about the stored graph; it is not a complete explanation of a model’s internal reasoning.

A bounded traversal example

This Cypher example illustrates a proposed schema: every node and relationship has a server-assigned organization_id, and every node has an allowed_users list. Identity parameters must come from authenticated server context, never from the model or user-supplied query text.

MATCH p = (start:Entity {id: $entity_id})-[:DEPENDS_ON*1..3]->(target)
WHERE all(n IN nodes(p)
          WHERE n.organization_id = $organization_id
            AND $user_id IN coalesce(n.allowed_users, []))
  AND all(edge IN relationships(p)
          WHERE edge.organization_id = $organization_id)
RETURN target.id AS target_id,
       [n IN nodes(p) | n.id] AS node_ids
LIMIT 50

The restriction checks every reached node and relationship, not only the starting node. Missing properties fail the predicate rather than granting access. The path is bounded; the query does not treat a variable-length relationship list as a single relationship with a scalar weight. See Cypher’s all() predicate.

This is a query illustration, not a security certification. It assumes the stated ACL schema, trustworthy ingestion, read-only credentials, and a server-controlled query template. Test cross-tenant edges, missing ACLs, revoked access, duplicate entity names, and restricted intermediate nodes. Query limits do not replace execution timeouts or database-level controls. Recheck authorization when fetching source passages and serving cached answers. The query has not been exercised against a live Neo4j deployment in this article.

Building and updating the graph

A reference pipeline parses sources, extracts candidate entities and relationships, resolves identities, records provenance, and indexes both graph and text. At question time, it selects a retrieval strategy, traverses permitted relationships, fetches supporting passages, and asks the generator to answer within that evidence.

Extraction quality is a first-class dependency. A generated relationship should not silently become a verified fact. Updates need to invalidate stale edges and summaries as well as old document chunks. Entity resolution, temporal validity, and permission propagation often require more work than the traversal itself.

When the complexity earns its place

Compare against lexical retrieval, vector retrieval, hybrid retrieval, and iterative retrieval on the same held-out questions. Include single-hop lookups as well as multi-hop and whole-corpus questions. Measure answer support, missing evidence, extraction errors, and permission failures, alongside build cost, update cost, and query latency.

A graph may help when relationships recur across questions and can be extracted or curated dependably. It may add little for straightforward passage lookup. A small corpus may be easier to search directly. The choice is workload-dependent.

Dated follow-up: small-scale alternatives

The earlier version included 2026 material under a November 2025 publication date. This revision preserves that original date and records the correction on September 14, 2026; it does not infer when the intermediate appendix was first added.

Microsoft’s GraphRAG documentation describes a particular indexing and querying system; it is not interchangeable with every graph-assisted retrieval design. LightRAG, HippoRAG, and Graphiti explore different retrieval or memory patterns. Compare their documented data models, update behavior, and operational requirements against your task. Repository popularity is not a deployment validation result.

The previous cross-project cost multipliers and release-status assertions have been removed because they did not provide a consistent, reproducible comparison. For a personal knowledge base, measure a small sample with a chosen version and actual token and infrastructure usage before extrapolating.

I would start with the question that simpler retrieval cannot answer reliably. If an explicit relationship fixes that failure, a graph has earned a place. If it does not, the diagram becoming more sophisticated is not a reason to keep it.

Keep reading

AI Search

Hybrid RAG

Combine lexical and vector rankings while preserving source content; an executable fusion example and explicit evaluation boundaries.

Read post