Similar is not the same as relevant
When an agent keeps getting your own documents wrong, the model is rarely the problem. Why search by similarity finds the pointer and never follows it - and what writing the connections down actually buys you.
By Joe WOct 7, 202611 min read
The most common complaint about AI systems inside a business isn't that they hallucinate. It's quieter than that: it keeps getting our own documents wrong.
The answer sounds plausible. It cites a real document. It just isn't the right one — and the person asking usually knows exactly which document it should have found.
This almost always has the same cause, and it isn't the model.
Two kinds of librarian
Imagine asking a librarian: "what are our payment terms with this supplier?"
The first librarian is extraordinarily well read. They've absorbed the language of every book in the building, so they bring you six books that talk about payment, invoicing and terms. All six are genuinely about your subject. None of them answers your question.
The second librarian knows something different. They know that your contract's payment clause points at Schedule B, that Schedule B is where "Net 30" is actually defined, and that there's an amendment from March that changed it. They bring you three documents, and together they answer the question.
The first librarian is how most AI retrieval works today. It's called vector search, and what it does is find text that sounds similar to your question.
Similar is not the same as relevant.
Where sounding similar isn't enough
Two failures show up constantly, and both are invisible in a demo with ten documents.1System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 6 (why naive RAG falls short, knowledge graphs, the hybrid RAG approach, retrieval optimisation) Open source (opens in a new tab)
The answer lives somewhere else. The clause you matched doesn't contain the answer — it points at the thing that does. Contracts do this constantly, and so do technical documents, policies, regulations and support tickets. Search by similarity and you reliably find the pointer and never follow it.
You match the heading and miss the detail. Long documents are hierarchies: a title, sections, subsections. Ask a question and the high-level summary at the top often looks like the best match, because it's phrased in exactly the general language you used. The actual answer is five sections down, in wording that doesn't resemble your question at all.1System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 6 (why naive RAG falls short, knowledge graphs, the hybrid RAG approach, retrieval optimisation) Open source (opens in a new tab)
Both failures have the same shape. The system found something related and had no way of knowing what that thing was connected to.
Writing the connections down
The fix is to stop treating your documents as a pile of loose text and start recording how they relate to each other — a knowledge graph.
That sounds grander than it is. A knowledge graph is a list of statements of the form this thing, this relationship, that thing:
- Clause 7.2 is part of Section 7
- Clause 7.2 references Schedule B
- Schedule B defines Net 30
- The March amendment supersedes Clause 7.2
None of that is clever. It's exactly the knowledge the second librarian had, written down in a form a program can follow.
The two working together
The important part is that this doesn't replace vector search. It finishes it.1System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 6 (why naive RAG falls short, knowledge graphs, the hybrid RAG approach, retrieval optimisation) Open source (opens in a new tab)
Vector search is excellent at one job: taking a question in ordinary language and finding a sensible place to start, even when the words don't match. That's a genuinely hard problem and the graph is no good at it. So the two run in order:
- Vector search finds the entry points — the handful of passages that look most like the question.
- The graph walks outward from those, following the relationships: up to the parent section for context, across to the schedule being referenced, forward to the amendment that superseded it.
- The original text is fetched last, for just that assembled set, and handed to the model.
There's a performance reason this order matters, and it's easy to get backwards. Walking relationships is only cheap when you start from a few known points. Searching the whole graph would be slow; following four hops out from five passages is not.1System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 6 (why naive RAG falls short, knowledge graphs, the hybrid RAG approach, retrieval optimisation) Open source (opens in a new tab) Vector search narrows, the graph enriches.
Where the connections come from
Somebody has to write all this down, and the answer is that a model does it — once, when the document arrives, not when somebody asks a question.
As each document is ingested it's split into sections, and a model reads each one with a narrow job: list the things this passage mentions, and what it points at. That output is stored as connections, the original text stays in ordinary storage, and the search index is built over the top.
// What the ingestion pass writes down for one chunk of a contract. The model
// reads the text once, offline, and records what it points at — not a summary,
// just the connections, so retrieval can follow them later.
{
"chunk_id": "c_8f21",
"document": "msa-2026-acme",
"entities": ["Payment Terms", "Schedule B", "Net 30"],
"relations": [
{ "from": "Clause 7.2", "type": "IS_PART_OF", "to": "Section 7" },
{ "from": "Clause 7.2", "type": "REFERENCES", "to": "Schedule B" },
{ "from": "Schedule B", "type": "DEFINES", "to": "Net 30" }
]
}This is the same move as doing the work before anyone asks. Reading every document carefully is far too slow to do while someone waits. Done at ingestion, it costs nothing at question time — you're just following links that were mapped weeks ago.
When you shouldn't build this
It would be dishonest to present this as a straight upgrade. It's real machinery: an extra store to run, an extraction step that can be wrong, and a schema of relationship types that needs maintaining as your documents change.
It pays for itself only when your documents genuinely point at each other. If your corpus is a few hundred help articles, each answering one question on its own, plain vector search is the right answer and a graph is a project you'll regret. Start where we did in wiring retrieval into an agent — keyword, semantic and a structured filter — and add this only when you can point at the failures.
Three symptoms say you've reached that point:
- Answers cite the right document but the wrong part of it.
- People keep saying "it missed the obvious one" — and the obvious one is always a document the matching one referred to.
- Your material is full of cross-references: contracts and their schedules, tickets and the defects that caused them, research notes and the filings underneath them.
That third one is why this matters for the systems we build. A deal room is not a pile of unrelated documents; it's a web of agreements, amendments and schedules that only make sense in relation to each other. The same is true of an investment file. The connections are the content.
The wider point
When an agent answers badly from your own documents, the instinct is to reach for a better model or a cleverer prompt. Usually neither is the problem. The model was handed six passages that sounded right and none that answered the question, and no model recovers from that.
Retrieval quality is a data modelling problem wearing an AI costume. How you split documents, what you record about them, and whether you keep track of how they relate — those decisions set the ceiling on every answer your system will ever give.
Sources
- System Design for the LLM Era: Patterns and principles for production-grade AI architecture (opens in a new tab)
Ch. 6 (why naive RAG falls short, knowledge graphs, the hybrid RAG approach, retrieval optimisation)
Thinking about an agent like this for your team?
Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.
