Skip to content
Back to Insights
ApplicationsBy KE Engineering Team

Streaming Interfaces: Frontend for AI

Streaming Interfaces: Frontend for AIApplications cover for Streaming Interfaces: Frontend for AItokensAPPLICATIONSStreaming Interfaces:Frontend for AI// FIRST TOKEN IS THE METRIC

A chat box that prints tokens is a weekend project. An interface that streams reliably to users on production networks isn't. The failure modes are all mundane and they all show up in production: someone switches tabs, a phone drops to one bar, a user hits back mid-generation, or the response is structured data and the UI renders half an object. If you haven't set a time-to-first-token budget, start with Latency Budgets for AI Features.

Pick SSE unless you have a reason not to

Server-Sent Events is the right default for this shape of problem. The traffic is one-directional, the payload is text, and the protocol gives you automatic reconnection and event ids for free. WebSockets are correct when you need genuine bidirectional communication, like live collaboration or voice, and they cost more to operate well.

Plain fetch streams work too and are simpler in some stacks, but you give up the reconnection semantics that make the next section possible.

Nobody builds resumability until it bites them

The naive implementation ties the generation to the HTTP connection. Connection drops, generation is lost, the user sees a truncated answer and retries, and you pay for the whole thing twice.

Decouple them. When a generation starts, give it an id and write tokens into a short-lived server-side buffer as they are produced. The client reads from that buffer. If the connection drops, the client reconnects with the last event id it received and picks up where it stopped. The model never knows anything happened.

Resumable stream topologyThe model writes into a stream buffer, the client reads from the buffer, and on reconnect the client replays from last-event-id before following the live generation.STREAM RESUMECLIENTEventSourceSTREAM BUFFERTTL, size capMODELgenerationtokensreconnect: last-event-id// THE CLIENT RECONNECTS, THE BUFFER REPLAYS
Fig. 1: The connection and the generation have different lifetimes.
typescript
export async function GET(req: Request) {  const url = new URL(req.url);  const streamId = url.searchParams.get("streamId")!;  const lastEventId = Number(req.headers.get("last-event-id") ?? -1);  const encoder = new TextEncoder();   const stream = new ReadableStream({    async start(controller) {      const send = (chunk: Chunk) =>        controller.enqueue(          encoder.encode(`id: ${chunk.seq}\ndata: ${JSON.stringify(chunk)}\n\n`)        );       // 1. replay whatever the client missed      for await (const chunk of buffer.readFrom(streamId, lastEventId + 1)) {        send(chunk);      }       // 2. then follow the live generation      const unsubscribe = buffer.subscribe(streamId, (chunk) => {        send(chunk);        if (chunk.done) {          unsubscribe();          controller.close();        }      });       req.signal.addEventListener("abort", () => {        unsubscribe();        controller.close();      });    },  });   return new Response(stream, {    headers: {      "Content-Type": "text/event-stream",      "Cache-Control": "no-cache",      "X-Accel-Buffering": "no",    },  });}

The buffer needs a TTL and a size cap. It's a transient replay log.

Your proxy is buffering your stream

That last header deserves its own section, because it's the most common cause of "streaming works locally but not in production."

A reverse proxy or CDN in front of your app will happily buffer the response and flush it in one piece, turning your carefully built stream into a slow request that arrives all at once. Nginx buffers proxied responses by default. Set X-Accel-Buffering: no, confirm compression is disabled for the event stream, and test through the full production path, beyond the app server. Everything works on localhost because there's no proxy in the way.

Cancellation costs money

When a user navigates away, the browser closes the connection. If nothing propagates that upstream, the model keeps generating and you keep paying. Wire the client abort through to the provider request, and treat an orphaned generation as a bug worth alerting on. Every token generated after the user leaves is still billed, so an orphaned generation costs as much as a finished one. LLM Cost Control Engineering covers the rest of the bill.

Rendering structured output as it arrives

If the response is JSON rather than prose, you can neither wait for all of it nor render half of it. A partially filled object shows a field with a truncated value, which reads as wrong while loading.

Parse incrementally, but only commit a field to the UI once it's complete. A partial JSON parser that reports which keys are finished and which are still arriving is worth the dependency. Render the finished fields, show a placeholder for the one in flight, and never display a string value that is still growing unless it's deliberately a prose field.

A strict schema pays off twice here. See Structured Output That Holds in Production. When the structured output is a component choice rather than data, Generative UI covers how to render it as it streams.

The accessibility trap

Putting aria-live="polite" on a streaming container seems right. In practice it makes most screen readers read fragments as they arrive, over and over, for the length of the generation.

Mark the streaming container aria-busy="true" while it fills and announce once on completion. Test with NVDA and VoiceOver, because behavior differs between screen readers and aria-busy support is uneven. Streaming is a visual affordance. For assistive technology the useful event isn't that tokens are arriving, it's that the answer is ready.