All articlesAI Engineering

Two kinds of question your agent can't tell apart

"What are our payment terms?" has an answer sitting in a document. "What themes run through all four hundred contracts?" does not - and your system will answer both with equal confidence.

Oct 8, 202611 min read

Two kinds of question your agent can't tell apart

Ask an AI system "what are our payment terms with this supplier?" and it does reasonably well.

Ask it "what themes run through all four hundred of these contracts?" and you get something that sounds thoughtful and is, on inspection, almost entirely invented.

Those two questions look similar. They are not the same kind of question at all, and the difference explains a failure that no amount of better search will fix.

Where the answer lives

A local question has its answer sitting somewhere specific. The payment terms are in a clause, in a document, on a page. The system's job is to find that passage and repeat it accurately. If retrieval is good, the answer is good — which is why most of the effort in these systems goes into retrieval, and why following the relationships between documents pays off.

A global question has no such passage. "What themes run through these?" is not written down in any of them. Nor is "how has our position changed over three years," or "what do our customers complain about most." The answer has to be constructed from the whole collection, and there is nothing to retrieve.

This is the part that catches teams out, and it's a known limit of ordinary retrieval rather than a bug in anyone's implementation: similarity search is built to find passages, so when no passage holds the answer it returns the closest ones anyway. The system doesn't refuse. It fetches the handful of passages that look most like the question, writes a confident summary of those, and presents it as a summary of everything. The output is a summary of five documents wearing the costume of a summary of four hundred.

The demonstration

Microsoft Research published the clearest illustration of this in early 2024, in the work that gave the technique its name.1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab)

Working with thousands of news articles, they asked a conventional system: "What are the top 5 themes in the data?" It produced generic filler — things like "improving the quality of life in cities" — which was not what the corpus was about at all.1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab)

A second example is sharper still. They asked about an entity that appeared throughout the collection. The conventional system answered: "The text does not provide specific information on what Novorossiya has done." It had retrieved passages that didn't happen to contain the word, and concluded the corpus was silent — when in fact the information was spread across many documents, none of which was individually the best match.1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab)

That's the failure in miniature. Not a wrong answer. A confident denial that the information exists.

The fix is to answer the question before it's asked

You can't retrieve an answer that isn't written anywhere. So you write it down in advance.

The approach: have a model read the entire collection and extract the entities and the relationships between them, building a map of the material. Then cluster that map into groups of closely-related things — bottom-up, so small clusters sit inside larger ones. Then summarise each cluster ahead of time. When a global question arrives, the system answers from those prepared summaries rather than from retrieved passages.1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab)

The corpus, in effect, describes itself in advance, at several levels of zoom.

If that sounds familiar, it should. It's the same move as doing the work before anyone asks: reading four hundred contracts is impossibly slow while somebody waits, and perfectly reasonable overnight. The difference is that here precomputation doesn't just buy speed — it's the only way the answer exists at all.

What the results actually showed

This is where it's worth being careful, because the technique is usually described more enthusiastically than its own authors described it.

The reported wins were in comprehensiveness and in giving the reader supporting material to check — answers that covered the ground and showed their sources.1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab) Those are real gains, especially the second one.

But on faithfulness — whether the answer sticks to what the source material actually says — they report that it "achieves a similar level of faithfulness to baseline RAG."1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab) Not better. Similar.

That deserves emphasis, because the pitch people repeat is that this technique reduces hallucination. Its own evaluation doesn't claim that. It claims more complete answers with better provenance, which is a different and more modest thing.

Two further caveats worth stating plainly. The evaluation leaned on a model acting as judge, scoring qualities like comprehensiveness and diversity1Jonathan Larson and Steven Truitt, Microsoft Research · Feb 13, 2024GraphRAG: Unlocking LLM discovery on narrative private data Open source (opens in a new tab) — and we've written about where judge models are the wrong instrument. And the published work says very little about cost, which is the first question a practitioner asks: having a model read an entire corpus, extract every entity, and summarise every cluster is not a cheap background job, and it has to be redone as the corpus changes.

How to tell which problem you have

Before building any of this, work out whether your users actually ask global questions. Most don't.

Mostly local: looking things up, checking terms, finding the clause, answering from a policy. Invest in retrieval quality. Nothing here applies to you, and building it would be a large bill for a capability nobody uses.

Genuinely global: "summarise the position across these," "what's changed," "what patterns are in this." Note that these questions usually come from a different person — an analyst or a partner rather than someone doing daily work, which is why they often surface late, after the system is built.

A useful test: ask whether a competent human could answer the question by reading three documents, or whether they'd have to read all of them. The first is a retrieval problem. The second isn't.

The honest middle path

For most of the systems we build, the sensible answer isn't a full corpus-wide graph with summaries at every level. It's narrower:

  • Be precise about which global questions matter — usually two or three, not an open set.
  • Precompute those specific summaries on a schedule, scoped to a deal, a portfolio or a client rather than everything you hold.
  • Make it obvious in the interface which kind of answer the user is getting, so a summary of a collection is never mistaken for a quote from a document.
  • Keep the ordinary retrieval path for everything else, because that's what most questions are.

The real lesson isn't that every system needs a knowledge graph. It's that your system is quietly answering two different kinds of question with one mechanism, and only one of them is being answered properly. Knowing which is which is most of the work.

Sources

  1. Jonathan Larson and Steven Truitt, Microsoft ResearchFeb 13, 2024

    GraphRAG: Unlocking LLM discovery on narrative private data (opens in a new tab)

Thinking about an agent like this for your team?

Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.