All articlesArchitecture

Fast AI products do the work before you ask

You can't make a language model think faster. You can stop it thinking while somebody waits. The precomputation trick behind every AI product that feels instant, in plain English.

Sep 29, 202611 min read

Fast AI products do the work before you ask

Ask a language model a real question and it takes a few seconds to answer. Sometimes ten. That is simply how long the thinking takes, and no amount of clever engineering on your side makes the model itself faster.

Yet some AI products feel instant. Not fast — instant, in a way that shouldn't be possible given what's happening underneath.

They aren't using a secret model. They've moved the slow part to a moment when nobody was waiting.

The restaurant trick

A good restaurant can put a complicated dish in front of you twelve minutes after you order it. The dish takes four hours to make.

This isn't a contradiction, because most of those four hours happened at nine in the morning. The stock was simmered, the sauces were reduced, the vegetables were prepped. At eight in the evening the kitchen is assembling, not cooking from scratch. Nobody sits at a table while a stock reduces.

That's the whole idea, and it's the single most effective thing you can do about AI latency. Look at every slow step in your system and ask one question: does this have to happen while someone is waiting?

Usually, a surprising amount of it doesn't.

AHEAD OF TIME — NOBODY WAITINGthe slow, expensive thinkingruns on a schedule, in the backgrounda shelf of ready answersLIVE — SOMEONE IS WAITINGa requestexpects an answer nowtake one — instant
The request never triggers the expensive work. It collects the result of expensive work that already finished.

Three products, one trick, three shapes

A recent book on designing systems around language models walks through four production architectures in detail.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers) Open source (opens in a new tab) What's striking, reading them side by side, is that three of them solve their hardest problem the same way — and the way looks completely different each time.

The code editor: don't redo what hasn't changed

An AI code editor has to understand your whole codebase to be useful, and understanding a codebase is slow. Doing it on every keystroke is impossible.

So it's done once, up front. The hard part is keeping that understanding current as you type without redoing it. The trick described in the book is a neat one: both your machine and the server keep a kind of nested fingerprint of the code — a short code for each file, which combine into a code for each folder, which combine into one code for the whole project.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers) Open source (opens in a new tab)

Think of two people with identical bookshelves in different cities, trying to find the one book that differs. They don't read out every title. They compare a single summary number for the whole shelf. If it matches, they're done. If it doesn't, they compare the summary for each row, then each book in the row that disagrees. A handful of questions instead of thousands.

The result: you change one function, and only that function is reprocessed. Everything else was already done, hours or days ago.

The learning app: hand out what's already made

A language-learning app wants each lesson personalised to the learner. Personalising with a model takes seconds. Doing it when the learner taps "next" would put a spinner between every exercise.

Instead, a background job picks and orders the next twenty or thirty exercises for each active learner and puts that list in a fast store. When the learner taps next, the app takes the top item off the list and serves it. No model is involved in that moment at all.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers) Open source (opens in a new tab)

The lesson isn't generated when you ask for it. It was waiting for you.

The shopping search: precompute the question, not the answer

This one contains the subtlest idea in the book, and it's worth slowing down for.

A shopping site runs a nightly job over yesterday's searches. For each common search, a model works out what people actually meant — that "organic aged basmati" is a category plus two attributes, and that someone searching it might also want brown rice or millet. That's genuinely expensive thinking, and it happens overnight.

But here's what gets saved. Not the list of products. The database query that finds them.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers) Open source (opens in a new tab)

Why it matters: prices change, things sell out. A saved list of products goes stale the moment stock moves, and you'd be showing people things they can't buy. A saved query re-runs against live data every time, so it's always current — you've kept the expensive part (working out what the person meant) and thrown away the part that rots (what was in stock last night).

Cache the recipe, not the cooked meal. The recipe is what took the thinking. The meal goes cold.

What happens to the first person?

There's an obvious hole in all of this. Precomputing works for people you knew about. What about someone brand new, or someone back after three months, with nothing waiting for them?

The answer is the part most teams skip, and it's what makes the pattern safe to ship: when the shelf is empty, don't make them wait for the model. Serve something reasonable using ordinary code — a simple rule, a sensible default, the most popular option for someone at their level — and kick off the background job at the same time. Their first answer is slightly less clever. Every answer after that comes off the shelf.1Sampriti Mitra, Packt Publishing · Jun 1, 2026System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers) Open source (opens in a new tab)

goserving/next.go
// Serving is a queue pop, not a model call. The expensive thinking already
// happened, in a background job nobody was waiting for.
func (s *Service) Next(ctx context.Context, userID string) (Item, error) {
    id, err := s.buffer.Pop(ctx, userID) // one ready item, removed as it is taken
    if err == nil {
        return s.items.Get(ctx, id)
    }
    if !errors.Is(err, ErrEmpty) {
        return Item{}, fmt.Errorf("read buffer: %w", err)
    }

    // Nothing ready: this is a new or returning user. Serve something decent
    // with ordinary code, then start the refill. Nobody waits for the model.
    item, err := s.rulesFallback(ctx, userID)
    if err != nil {
        return Item{}, fmt.Errorf("fallback: %w", err)
    }
    s.refill.Trigger(ctx, userID)
    return item, nil
}

That is the whole shape, and it's about fifteen lines. The interesting engineering isn't here — it's in the background job that keeps the shelf stocked.

What this looks like in our own agents

We reach for the same move constantly, though it rarely looks as tidy as a book diagram.

Retrieval is the clearest case. When one of our agents answers a question about a client's documents, nothing about those documents is being read or understood at that moment. That work happened at ingestion: the documents were split into sections, turned into a searchable form, and indexed. The live request does a lookup, not an analysis — which is exactly the split we described in wiring retrieval into an agent.

Turning text into its searchable form is itself slow enough to be worth never repeating, so those results are kept and reused. The same piece of text, seen twice, is only ever processed once.

And for genuinely long work — producing a document, running a multi-step task — we don't pretend it's instant. The request returns immediately with a task to follow, and results stream back as they're produced, which is the mechanism behind tasks and push in A2A. It isn't precomputation, but it comes from the same instinct: never leave a person watching a blank screen while a model thinks.

Where it doesn't work is just as worth saying. You can only precompute an answer if you can guess the question. Our agents answer things nobody could have predicted — a specific question about a specific clause in a specific contract — and for those there is no shelf to stock. That work is slow because it has to be.

Three questions to ask about your own system

If you take one thing from this, take these:

  • Does this have to happen now? Indexing, summarising, categorising and enriching almost never do. They can happen when the data arrives instead of when someone asks about it.
  • Are we redoing work that hasn't changed? If one document in ten thousand was edited, only that one should be reprocessed. Getting this right is usually a bigger win than anything you'll do to the model.
  • Are we saving the thing that lasts, or the thing that rots? Save the interpretation, the structure, the query. Be careful about saving the answer, because the world underneath it moves.

The part nobody tells you

This all has a cost, and it isn't money. It's that you now have two systems instead of one: the thing serving people, and the thing filling the shelf. They can disagree. The shelf can run empty. The background job can quietly fail on a Sunday and nobody notices until answers get worse.

So the honest version of this advice is: move the slow work off the request path, and then watch the machinery that does it at least as closely as you watch the part users touch. The queue depth, the refill rate, the age of what's on the shelf. When an AI product that felt instant suddenly feels sluggish, it's almost never the model. It's the shelf, empty, and everyone falling through to the slow path at once.

Sources

  1. Sampriti Mitra, Packt PublishingJun 1, 2026

    System Design for the LLM Era: Patterns and principles for production-grade AI architecture (opens in a new tab)

    Ch. 2 (caching strategies, hybrid processing); Ch. 3 (Merkle tree sync); Ch. 4 (proactive curation, warm and cold paths); Ch. 5 (offline query pipeline, cache tiers)

Thinking about an agent like this for your team?

Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.