Cost, speed, quality. You get two.
Every AI system trades one of the three against the other two, and the expensive mistake is refusing to choose. How to decide which corner you are giving up - starting from who is waiting for the answer.
By George OnyangoOct 7, 202610 min read
Anyone who has built a house knows the sign on the contractor's wall: fast, good, cheap — pick two. It's a joke that survives because it's true, and because saying it out loud forces a conversation that would otherwise happen too late.
AI systems have their own version of that triangle. Most teams meet it without ever naming it, which is why so many AI projects end up in the worst possible place: expensive and slow, with quality nobody is happy with either.
This is the conversation worth having at the start.
The three corners
The clearest statement of this we've read puts it like this: cost is driven mainly by how much you send the model, speed by how much it writes back and how long you spend gathering context first, and quality by how much relevant context the model gets and how capable that model is. You can optimise two of those. The third moves against you.1System Design for the LLM Era: Patterns and principles for production-grade AI architectureCh. 1 (the cost, latency and quality constraints); Ch. 2 (model routing, caching, hybrid processing) Open source (opens in a new tab)
In plainer terms:
- Cost is mostly what you feed in. Every document you attach, every piece of history, every instruction — multiplied by every request, every day. It's the one that's invisible until the invoice arrives.
- Speed is mostly what comes out, plus everything you did before the model started writing. Longer answers take longer. So does searching your documents first.
- Quality is mostly about context and capability. The more relevant material the model has in front of it, and the stronger the model, the better the answer.
Why they pull against each other
Every lever you have improves one corner at the expense of another. There are no exceptions, which is what makes it a triangle rather than a list of good ideas.
- Use a stronger model. Quality up. Cost up, speed down.
- Give it more context — more documents, more history. Quality up. Cost up, speed down.
- Use a smaller, faster model. Cost down, speed up. Quality down.
- Cache and precompute. Speed up, cost down — at the risk of serving something slightly out of date.
- Ask for shorter answers. Speed up, cost down. Completeness down.
Read that list again and you'll notice something. Almost every instinct that makes an AI feature better makes it slower and more expensive, and almost every instinct that makes it cheaper and faster makes it worse. That's not a flaw in anyone's engineering. It's the shape of the problem.
The expensive mistake is refusing to choose
Here's how it usually goes.
The first version is a bit vague, so you add retrieval to ground it in real documents. Better — and now slower, because you search before you answer. So you reach for a faster model, and quality dips, so you move back to the strong one and add caching to recover the speed. The cache goes stale, so you add invalidation. Six months later the system is slow, expensive and complicated, and nobody can say which decision caused it.
Every individual step was reasonable. What was missing at the start was a sentence saying which corner the product was willing to give up.
You don't avoid the trade-off by not discussing it. You just make it by accident, repeatedly, in small decisions nobody wrote down.
How to choose: start with who is waiting
The decision gets much easier when you stop asking "how good can this be?" and start asking three questions about the job itself.
Is a person sitting there waiting? If someone is watching a cursor blink, speed is the corner you protect. A good answer in eight seconds loses to a decent answer in one, because the slow one doesn't get used. In-product assistants, support, search — speed first.
What does being wrong cost? If a wrong answer goes into a document a client acts on, quality is the corner you protect, and you simply remove the expectation of speed. Run the work as a job, show progress, deliver in minutes. Research, analysis, anything that becomes a deliverable — quality first.
How many times a day does this run? If the answer is tens of thousands, cost is the corner you protect, and quality per item matters less than consistency across all of them. Classifying, tagging, enriching, extracting — cost first.
Notice that none of those questions is about the model. They're about the job, which is why this is a product decision that happens to have an engineering consequence, not the other way round.
Different corners for different jobs, in the same product
The triangle applies per task, not per company. A single product usually needs all three answers.
The in-app assistant protects speed. The overnight analysis protects quality. The ingestion pipeline that reads every incoming document protects cost. Three different trade-offs, three different model choices, one product — which is exactly why we keep the model behind a thin boundary rather than picking one and standardising on it everywhere, as we wrote in swapping the brain.
Some of the trade also disappears if you move work off the request path entirely. Precomputing doesn't break the triangle, but it does let you buy speed with infrastructure instead of with quality — the pattern in doing the work before anyone asks.
Make the trade visible, or you'll optimise the wrong thing
One practical warning. Teams reliably optimise whichever corner they can actually see, and in most organisations that's cost, because it arrives as a bill with a number on it. Speed shows up as complaints. Quality shows up as nothing at all, until a client notices.
So measure all three, deliberately:
- Cost per completed task, not cost per request. A cheap model that needs three attempts isn't cheap — the metering approach we described in what an agent actually costs.
- Speed at the slow end, not on average. Averages hide the experience that makes people stop using a feature.
- Quality against a fixed set of real tasks, scored the same way every time, so a change that trades quality for speed shows up as a number rather than a feeling — with arithmetic wherever the failure has a measurable shape.
What we tell clients
When we scope an agent, one of the first things we put in writing is which corner we're trading and why. Not because it's a formality, but because it's the decision that everything else follows from — the model, the retrieval design, whether the work happens live or in the background, what the interface promises.
A system that answers in under a second, costs very little to run, and reasons deeply over your entire document history does not exist. A system that answers in under a second, costs very little, and handles the eighty per cent of questions people actually ask — that exists, and it's usually the right thing to build first.
Pick two, on purpose, and say which one you gave up.
Sources
- System Design for the LLM Era: Patterns and principles for production-grade AI architecture (opens in a new tab)
Ch. 1 (the cost, latency and quality constraints); Ch. 2 (model routing, caching, hybrid processing)
Thinking about an agent like this for your team?
Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.