Skip to content
Back to Insights
AI EngineeringBy KE Engineering Team

Context Engineering: What Goes in the Window

Context Engineering: What Goes in the WindowAI Engineering cover for Context Engineering: What Goes in the Windowtool defs + systemtaskcontext budget consumedAI ENGINEERINGContext Engineering:What Goes in theWindow// WHAT EARNS THE WINDOW

The context window is a budget. Every token in it either helps the model do the task or gets in the way. Most production systems fill it carelessly: the full system prompt, every tool definition, the entire conversation, whatever retrieval returned, in whatever order. Then they wonder why quality degrades on long sessions.

Context engineering is the discipline of deciding what goes in, in what order, at what fidelity. It matters more than prompt wording, and most teams never budget for it.

Order matters

Models attend unevenly across the window. Content at the beginning and end gets more weight than content in the middle. The effect is documented in Lost in the Middle: How Language Models Use Long Contexts (2023), and while newer long-context models have narrowed it, the practical consequence stands: decide the order on purpose.

Put the most important instructions first and the most important data last, immediately before the question. Put reference material in the middle, where lower attention costs the least. Tool definitions usually sit where the provider's API places them, so this applies to what you assemble yourself. If retrieval returned eight documents, the most relevant one goes at the end.

Allocate the budget explicitly

Treat the window as a set of named allocations with hard caps.

typescript
const budget = {  system: 0.10,      // fixed  tools: 0.12,       // curated per task, never the full catalog  memory: 0.05,      // retrieved facts, ranked  retrieval: 0.30,   // ranked, deduplicated, truncated  history: 0.23,     // compacted, most recent verbatim  reserve: 0.20,     // headroom for the response}; // Fractions of the usable window, so a model swap doesn't move the caps. Each section is assembled to fit its cap.

When a section exceeds its cap, it is compacted or truncated according to its own rule. History gets summarized. Retrieval gets ranked and cut. Tools get filtered to the ones relevant to the task. The alternative is that one section crowds out another and you find out when the model forgets an instruction that was pushed out of the window.

Compaction is lossy; decide what to lose

Long conversations have to be compressed. The naive approach is to summarize the whole history when it gets long. The better approach is tiered fidelity: the most recent turns stay verbatim, older turns become a structured summary, and the oldest become a few retained facts.

What to preserve through compaction: decisions that were made, constraints the user stated, entities that were named, and anything the model committed to. What to drop: pleasantries, superseded plans, and intermediate reasoning that led nowhere.

Summarize with a small model and a fixed schema, so compaction is cheap and its output is predictable. A summary that is itself free-form prose will drift.

Retrieval returns candidates, not context

The top-k results from a vector search are candidates. They aren't ready to be placed in the window.

Deduplicate near-identical chunks. Rerank with a cross-encoder or a small model, because embedding similarity is a coarse first pass. Truncate each chunk to the relevant part of the passage. Then place them in relevance order with the best last. A retrieval step that dumps eight raw chunks into the window is spending a third of the budget to communicate one document's worth of information.

Tool definitions are context too

Every tool in the catalog costs tokens on every request whether or not it is used, so curate per task; Tool Design Is Prompt Design covers how.

Measure it

Instrument context composition per request: tokens per section, what was truncated, what was compacted. When quality degrades on long sessions, this telemetry is what tells you whether the instruction got crowded out or the retrieval got noisy. Without it you are guessing.

Make the budget a test

A budget only holds if something enforces it. Express each allocation as a fraction of the usable window rather than a fixed token count, so a model swap doesn't break it, and assert on it in the eval suite the way you would assert on a type.

typescript
test("tool schemas stay inside their allocation", async () => {  const run = await agent.run(fixtures.multiStep);  expect(overBudget(run.segments, run.window)).toEqual([]);});

Enforce in CI, alert in production. Failing a user request because a schema grew is worse than the degradation you're guarding against.

The discipline

Prompt engineering asks how to say it. Context engineering asks what deserves to be said at all, given a fixed budget and a model that reads unevenly.