All articlesAI Engineering

The clock your user actually feels

Two systems answer in nine seconds. One of them gets reported as broken. Why time to first token is a different measurement from total time - and why streaming can hide a system that is quietly getting slower.

Oct 2, 20269 min read

The clock your user actually feels

Two systems answer the same question in exactly nine seconds.

The first shows nothing for nine seconds, then the whole answer at once. People reload the page, click the button again, and tell you it's broken.

The second starts writing after half a second and keeps going for the rest. People read along, and nobody mentions the speed at all.

Same nine seconds. One of them has a performance problem and it isn't the slow one, because there isn't a slow one.

Two different clocks

There are two measurements here and teams usually only watch one.

Total time is how long until the answer is complete. It's what your monitoring records and what appears in a report.

Time to first token is how long until something appears. It's what the person experiences as "the wait," and when it goes wrong it's the one that produces complaints.

A waiter who says "that'll be about twenty minutes" is doing the same job as one who walks away without speaking. Only one of them gets complained about.

Stream the work, not just the words

Streaming text as it's generated is the obvious move and most teams have it. For an agent, it isn't enough — because before any text exists, the agent is often doing the slowest part of its job.

An agent that searches your documents, reads four of them, and then writes might spend most of its time before the first word. Streaming the writing doesn't help, because the silence happens earlier.

So stream the steps as well:

  • Searching contracts…
  • Reading 4 documents…
  • Drafting…

This is not decoration. It converts an unexplained wait into visible progress, and it tells the person whether the system understood them — if it says "searching invoices" and you asked about contracts, you've learned something useful three seconds in rather than nine.

When the work is genuinely long, stop pretending

Streaming is the right answer for seconds. It is the wrong answer for minutes.

If a job takes four minutes, holding a connection open and trickling output is fragile — the connection drops, the user navigates away, the phone locks, and the work is lost or orphaned. Long-running work wants a different shape: accept the request, return something to follow, and push updates as they happen. That's the mechanism behind tasks and push notifications in A2A, and it means closing the laptop doesn't cancel the job.

The rule of thumb we use: seconds, stream it; minutes, make it a task with progress; and if it's only slow because of work that didn't need to happen now, it probably belongs off the request path entirely.

Buying the first token back

Two techniques worth knowing, in increasing order of effort.

Race the providers. If the primary model hasn't started responding within your budget, fire the same request at a faster backup and take whichever begins first. You pay for both on the requests where it triggers, which is why it's worth reserving for the paths where responsiveness genuinely matters.

Have a small model start while a large one thinks. A fast, cheap model can draft ahead and a stronger one verify in larger batches, which is how products achieve a first word that appears essentially instantly. This is real engineering rather than a configuration change, and it only pays back at serious volume.

Before either, check the obvious: how much of your time-to-first-token is retrieval, prompt assembly, and your own services? Frequently most of it is, and that's cheaper to fix than anything involving the model.

The trap: streaming hides a slow system

Here's the part that deserves a warning.

Streaming makes a slow system feel acceptable. That's the point — and it's also how a system that takes forty seconds to finish an answer survives for months without anyone escalating it, because every individual user watched words appear and assumed that was normal.

So watch both numbers, and alert on both. First token tells you whether the experience is acceptable. Total time tells you whether the system is actually healthy. A rising total with a flat first token is a system quietly degrading behind a comfortable façade.

One tension worth naming

Streaming and checking the output fight each other.

If you validate an answer before showing it — and we think you should, for failures with a measurable shape — then by definition you can't have shown it while it was being written. Anything you streamed is already on the screen when your checks run.

The resolution we use is to separate the two things an agent produces. Prose streams, because the cost of a bad sentence appearing briefly is low. Actions never stream — nothing that sends, pays, writes or deletes happens while text is still arriving. Those wait for the full output, the checks, and where it matters, a person.

Which gives a clean rule: stream what the user reads, gate what the system does.

Thinking about an agent like this for your team?

Describe the job you want automated and our Automation Architect blueprints it in seconds — or talk it through with the engineers who ship them.