Skip to content
Back to Insights
ApplicationsBy KE Engineering Team

Latency Budgets for AI Features

Latency Budgets for AI FeaturesApplications cover for Latency Budgets for AI Featuresp50p99first tokenAPPLICATIONSLatency Budgets for AIFeatures// MOMENTS, NOT ONE NUMBER

A traditional latency budget is a single number: this endpoint responds in under 300 milliseconds at p99. An AI feature doesn't have one latency. It has a sequence of moments, and users judge each one differently.

Treat the feature as a timeline and budget each point on it.

The four moments

Acknowledgment. The moment the user's input is visibly received. Budget: under 100ms. This is the same as any interface. If nothing changes on screen after they hit enter, they hit it again.

First token. The moment content begins appearing. Budget: under 1 second at p50, under 2 at p99. This is the number users experience as "how fast is it." Everything before this feels like waiting; everything after feels like reading.

Useful completion. The moment the user has enough to act on, which is usually before the generation finishes. For a structured response, this is when the key field arrives. Budget: depends on the task. For a drafted reply, a reasonable target is the first full sentence within 2 seconds. Instrument it separately from full completion.

Full completion. Generation ends. Users will wait here as long as the stream keeps moving. Budget: bounded, with a hard cutoff, but not the number to optimize first.

Where the time goes

user input  │ 20ms   client validation, acknowledge  │ 40ms   auth, rate limit, request routing  │ 180ms  retrieval (embed query, search, rerank)      ◀ often the surprise  │ 60ms   context assembly, prompt render  │ 600ms  model time-to-first-token                    ◀ dominant, variable  │ ...    streaming, ~40 tokens/s  ▼

That's 900ms before the first token. The budget is 1,000.

Retrieval is where budgets break. A vector search that was 30ms in development becomes 200ms with a reranker and a large index. It sits directly on the path to first token. Measure it in isolation, and run the steps that don't depend on its results (routing, first-turn classification) in parallel with it.

Spend the budget where users feel it

Optimizing full completion time by 20% is invisible if first token is slow. Optimizing first token by 300ms is felt by everyone. Rank work by which moment it moves.

Concretely: prompt caching skips reprocessing the cached prefix, so long repeated prompts reach the first token faster. A smaller model on the first-turn classification step cuts it further. Parallelizing retrieval with anything that doesn't depend on it cuts it again. None of these change total generation time much, and all of them change the experience.

Skeleton first, then substance

The interface should change state at every moment on the timeline. Acknowledge instantly. Show a structured skeleton at first token, shaped like the response that is coming. Fill in as it arrives. A spinner that persists until full completion throws away every intermediate signal the system produced.

Set the cutoff and enforce it

Give every AI call a hard timeout, and make it shorter than any proxy or gateway in front of it. A model call that can run for 60 seconds behind a 30-second gateway timeout produces an error the user sees and a generation you still pay for.

Decide what happens at the cutoff: return what has arrived, with an indicator that it was truncated. That beats a blank error, and it tells the user the truth.

Instrument every moment

Time-to-acknowledgment, time-to-first-token, time-to-useful, time-to-complete, each at p50 and p99, each broken down by the stage that consumed it. When the feature feels slow, this is what tells you whether retrieval regressed, the model got slower, or the front end stopped rendering skeletons.