When ten thousand people ask the same question at once
A cache answers questions it has seen before. During an incident it has seen none of them, and every request misses at the same moment. How to do the work once instead of four thousand times.
By Patrick WaweruOct 8, 202610 min read
Your biggest customer has an outage at 09:14 on a Tuesday. Within two minutes, four thousand of their staff open your product and type a version of the same question.
You have a cache. It does not help you at all.
Why the cache is useless at exactly the wrong moment
A cache answers questions it has seen before. At 09:14 it has never seen this one. So the first request misses — and so does the second, and the four-thousandth, because none of them finishes before the others start.
Every one of those requests goes to the model. You pay for four thousand identical answers, your rate limit trips, and the provider starts refusing you — at the exact moment your product most needed to work.
Picture a ticket office where everyone arrives at once and each person is sent individually to the back room to fetch the same ticket. The sensible clerk fetches one and photocopies it. That's the whole fix.
Collapse the duplicates
The pattern has a name — request collapsing, or coalescing. The gateway spots that a request already in flight is identical to the one arriving, holds the newcomer, waits for the first answer, and hands that single response to everyone waiting.
In Go this is almost embarrassingly easy, because the standard extended library ships it:
// One flight per distinct question. The first caller does the work; everyone
// who asks the same thing while it is still running waits for that result
// instead of starting their own.
var flights singleflight.Group
func (s *Service) Answer(ctx context.Context, q string) (Answer, error) {
key := normalise(q)
if cached, ok := s.cache.Get(ctx, key); ok {
return cached, nil
}
v, err, shared := flights.Do(key, func() (any, error) {
a, err := s.model.Answer(ctx, q) // the only call that reaches the provider
if err != nil {
return nil, fmt.Errorf("answer %q: %w", key, err)
}
s.cache.Set(ctx, key, a, ttl)
return a, nil
})
if err != nil {
return Answer{}, err
}
if shared {
metrics.CoalescedHits.Inc() // how many requests this one call served
}
return v.(Answer), nil
}Four thousand requests, one model call. The others wait roughly as long as the first one does, which is the same wait they'd have had anyway — except now you paid once.
The shared flag is worth keeping. It tells you how many requests each real call served, which is the number that proves this is doing anything. On a quiet day it will be 1. During an incident it will be startling.
The hard part: "identical" barely happens
That whole mechanism hangs on two requests having the same key. With natural language, they almost never do. Four thousand people describing one outage will produce four thousand different sentences.
So you need keys that are less literal. Three layers, cheapest first.
Normalise before hashing. Trim whitespace, lowercase, strip trailing punctuation. Crude, free, and it collapses a surprising number of near-duplicates that differ only in formatting. Be careful with anything cleverer — stripping words that look unimportant can quietly merge two questions that meant different things.
Match on meaning, not characters. Convert the question into its numeric form and look for a previously answered question that sits very close to it — in practice a similarity above about 0.95. "How do I reset my password" and "password reset not working" never match as text and match comfortably as meaning.
Cache the conversion too. Turning text into its numeric form is itself a paid call with its own latency. The same text, seen twice, should only be converted once. This one is pure profit and teams skip it constantly.
What exactly are you storing?
One decision decides whether this helps or hurts: do you cache the answer, or the work that produced it?
Cache the answer and you get the fastest possible response — and you will eventually serve something that was true an hour ago. Cache the retrieval and reasoning instead, re-running only the final step against live data, and the answer stays current at the cost of a little speed. We went through this trade in detail in doing the work before anyone asks: cache the recipe, not the cooked meal.
For an incident specifically, the answer cache is usually right. The situation is changing by the minute, so a short time-to-live — minutes, not hours — gets you the stampede protection without serving yesterday's advice tomorrow.
The cost angle, which is the easiest to sell
Request collapsing is one of the few optimisations that improves every corner at once: fewer calls, lower bill, less rate-limit pressure, and a system that stays up when it's most visible. That's unusual — most choices in these systems force a trade between cost, speed and quality.
It only shows up if you're measuring per-agent spend rather than one monthly total, because the saving is concentrated in spikes. If your bill is a single number at the end of the month, a four-thousand-request stampede looks like a rounding error — right up until the month it isn't, as we argued in what an agent actually costs.
Three ways to get this wrong
Collapsing requests that aren't actually the same. If your key ignores who is asking, two tenants with identical questions share an answer — and one of them gets the other's data. The key must include the identity and scope of the asker, always. This is the same failure as sharing state without an owner, arriving through a cache instead of a variable.
Caching failures. If the first call errors and you store that, you've just served the same error to four thousand people efficiently. Cache successes only.
Letting one slow call block everyone. Collapsing means every waiting request inherits the first one's fate. Put a timeout on the flight, so a hung provider call fails the group quickly rather than holding thousands of connections open.
The general shape
Most capacity problems in AI systems aren't about sustained load. They're about correlation — real users do the same thing at the same moment, because something happened in the world that made them.
Average throughput won't show you this. The time to look is the minute when everyone arrives at once, and the question to ask is whether your system does the work once or four thousand times.
Thinking about an agent like this for your team?
Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.