Skip to content
Back to Insights
AI EngineeringBy KE Engineering Team

LLM Cost Control Engineering

LLM Cost Control EngineeringAI Engineering cover for LLM Cost Control Engineering01 PROMPT CACHE02 SEMANTIC CACHE03 ROUTING$ / MONTHAI ENGINEERINGLLM Cost ControlEngineering// THREE LEVERS · IN ORDER

Your finance team is going to notice your LLM bill. Every team we work with hits this moment somewhere between month three and month six of production. The model providers keep charging what they charge, and prompt engineering alone doesn't fix it.

Three engineering levers move the cost number the most: prompt caching, semantic caching, and routing, in that order. Two more are worth doing once those are in place: batch APIs for offline work, and trimming what goes into the window, which Context Engineering covers.

Lever 1: Provider prompt caching (do this first)

For any workload with repeated context, prompt caching is the largest and easiest saving available: 50 to 90 percent off the repeated portion (figures as of mid-2026), with no model downgrade and no quality tradeoff. The model, the prompt, and the output stay exactly the same. You just stop paying full price to re-send the same prefix on every call. If you do nothing else on this list, do this one.

Anthropic, OpenAI, and Google all support prompt caching now, with similar economics: cache writes cost slightly more than a normal call, cache reads cost a fraction of a normal input call.

The pattern is simple: put the stable parts of your prompt (system instructions, tool definitions, retrieved context you'll reuse) at the beginning, mark them cacheable, put the volatile parts (user query, latest turn) at the end. The provider caches the prefix. Repeated calls only pay full price for the tail.

Prompt caching structureA prompt shown as three segments: a large cached prefix of system and tools, an optional cached assistant history, and a smaller uncached user query on the right.PROMPT CACHEone prompt, three segmentsSYSTEM + TOOLS + RETRIEVED CONTEXTcachedASSISTANTHISTORYUSER QUERYuncachedreads at a fraction of base input cost (varies by provider)optional boundaryfull input pricecache boundary → tokens billed at full rate →// STABLE PARTS FIRST, VOLATILE PARTS LAST
Fig. 1: Cache the prefix. Pay only for the tail.
ts
const response = await anthropic.messages.create({  model: MODEL,  // the snapshot you validated against  tools: [    ...TOOLS.slice(0, -1),    { ...TOOLS.at(-1), cache_control: { type: "ephemeral" } }  // caches every tool before it  ],  system: [    {      type: "text",      text: SYSTEM_INSTRUCTIONS,  // stable across requests      cache_control: { type: "ephemeral" }    }  ],  messages: [    { role: "user", content: userQuery }  // volatile, uncached  ]});

Two things to watch: cache TTLs (typically 5 minutes for ephemeral caches, expect misses if traffic is bursty) and the cache-hit-rate metric on your provider dashboard. If you're paying to write the cache and not reading it back, you're worse off than uncached. Instrument this from day one.

Lever 2: Semantic caching

The layer above provider caching. Instead of caching the prompt, cache the answer. When a new query comes in, embed it, look up nearest neighbors in a vector store, and if similarity is above a threshold, return the cached response.

Semantic caching flow with similarity gateA query is embedded and searched against a vector store. A similarity gate routes above-threshold matches to a cached answer, and misses to the LLM and back into the store, converging on a single response.SEMANTIC CACHEQUERYEMBEDVECTOR SEARCHSIMILARITY≥ τ ?default τ = 0.95 cosineYESCACHED ANSWERNOCALL LLMSTORERESPONSE// MISSES CALL THE LLM, THEN STORE THE ANSWER
Fig. 2: The answer cache gates on similarity.

Semantic caching pays off in workloads with repetitive user intent: customer support, internal Q&A over documentation, common analytical queries. It fails in workloads where every query is unique (creative generation, deep code work).

The most important knob is the similarity threshold. Too high and hit rate is negligible. Too low and you return stale or wrong answers. Start at 0.95 cosine similarity, measure the false-positive rate on a validation set, tune down carefully. Never automate this without a human review of the false-positive class.

Cache invalidation is what breaks first. If the underlying documents or data change, the cached answer is now wrong. Version your cache entries against a data version, invalidate when the source moves.

Key the cache on the tenant as well as the query, and on anything else that scopes the answer, such as role or region. A cache keyed on query text alone serves one customer another customer's answer; Multi-Tenant AI covers why that is the worst bug in the system.

Lever 3: Routing (model tiering)

Not every query needs the flagship model. A well-designed system routes cheap queries to a small model and reserves the expensive model for hard ones.

Two routing patterns work in production:

Explicit routing by task type. Classify at ingress. Simple summarization goes to the smallest model in a provider's current lineup. Complex reasoning goes to the mid-tier or flagship model. Code generation goes to whatever your team has evaluated best for code. Route based on a small classifier or on the endpoint the client called.

Cascade routing. Try the small model first. If confidence is low or the output fails validation, escalate to the larger model. Works well when the small model gets it right most of the time and the escalation path is a rare exception.

We usually recommend explicit routing over cascade. Cascade adds latency (two model calls when the small one fails) and complexity (now you need a confidence signal). Explicit is boring and predictable.

Evals decide how small you can go

Routing raises an uncomfortable question: how do you know the small model is good enough? Bigger models cost more and are usually slower, and plenty of tasks don't need frontier-level intelligence. Data extraction, classification, format conversion, and similar well-scoped work is often handled correctly by a smaller, cheaper, faster model.

You find out with evals. Run the candidate task through your eval set on the smaller model. If it clears your quality bar, route it there. If it doesn't, you now know exactly where it fails and by how much, which is a better argument than anyone's intuition about model tiers.

Structured data extraction is the strongest candidate to try first. The task is well-defined and the output is checkable, so the eval is cheap to write and the pass or fail is unambiguous.

We covered how to build that eval set in Evals for Production Agents.

The order matters

Prompt caching first. It's free savings on your current architecture. Semantic caching second, and only if your workload has repeated intent. Routing third, and only after you've measured which queries need the flagship model.

Teams often reverse this. They start by picking a cheaper model, then wonder why quality dropped. Cost engineering starts with keeping the same model and paying less for the same calls. It ends with using cheaper models only where the quality tradeoff is acceptable.

What to instrument

Cost per successful request, and total spend. Cache hit rate per prompt template. Router accuracy (is the small model good enough for the queries you're sending it?). Cost per user, per feature, per tenant if you're multi-tenant. When one customer's usage explodes, you need to know before your CFO does.

The bill is going to keep growing. The job is to make sure it grows because usage grew, not because each request got more expensive.