Context Engineering: What Goes in the Window
The context window is a budget. Every token in it either helps the model do the task or gets in the way. Most production systems fill it carelessly: the full system prompt, every tool definition, the entire conversation, whatever retrieval returned, in whatever order. Then they wonder why quality degrades on long sessions.
Context engineering is the discipline of deciding what goes in, in what order, at what fidelity. It matters more than prompt wording, and most teams never budget for it.
Order matters
Models attend unevenly across the window. Content at the beginning and end gets more weight than content in the middle. The effect is documented in Lost in the Middle: How Language Models Use Long Contexts (2023), and while newer long-context models have narrowed it, the practical consequence stands: decide the order on purpose.
Put the most important instructions first and the most important data last, immediately before the question. Put reference material in the middle, where lower attention costs the least. Tool definitions usually sit where the provider's API places them, so this applies to what you assemble yourself. If retrieval returned eight documents, the most relevant one goes at the end.
Allocate the budget explicitly
Treat the window as a set of named allocations with hard caps.
const budget = { system: 0.10, // fixed tools: 0.12, // curated per task, never the full catalog memory: 0.05, // retrieved facts, ranked retrieval: 0.30, // ranked, deduplicated, truncated history: 0.23, // compacted, most recent verbatim reserve: 0.20, // headroom for the response}; // Fractions of the usable window, so a model swap doesn't move the caps. Each section is assembled to fit its cap.When a section exceeds its cap, it is compacted or truncated according to its own rule. History gets summarized. Retrieval gets ranked and cut. Tools get filtered to the ones relevant to the task. The alternative is that one section crowds out another and you find out when the model forgets an instruction that was pushed out of the window.
Compaction is lossy; decide what to lose
Long conversations have to be compressed. The naive approach is to summarize the whole history when it gets long. The better approach is tiered fidelity: the most recent turns stay verbatim, older turns become a structured summary, and the oldest become a few retained facts.
What to preserve through compaction: decisions that were made, constraints the user stated, entities that were named, and anything the model committed to. What to drop: pleasantries, superseded plans, and intermediate reasoning that led nowhere.
Summarize with a small model and a fixed schema, so compaction is cheap and its output is predictable. A summary that is itself free-form prose will drift.
Retrieval returns candidates, not context
The top-k results from a vector search are candidates. They aren't ready to be placed in the window.
Deduplicate near-identical chunks. Rerank with a cross-encoder or a small model, because embedding similarity is a coarse first pass. Truncate each chunk to the relevant part of the passage. Then place them in relevance order with the best last. A retrieval step that dumps eight raw chunks into the window is spending a third of the budget to communicate one document's worth of information.
Tool definitions are context too
Every tool in the catalog costs tokens on every request whether or not it is used, so curate per task; Tool Design Is Prompt Design covers how.
Measure it
Instrument context composition per request: tokens per section, what was truncated, what was compacted. When quality degrades on long sessions, this telemetry is what tells you whether the instruction got crowded out or the retrieval got noisy. Without it you are guessing.
Make the budget a test
A budget only holds if something enforces it. Express each allocation as a fraction of the usable window rather than a fixed token count, so a model swap doesn't break it, and assert on it in the eval suite the way you would assert on a type.
test("tool schemas stay inside their allocation", async () => { const run = await agent.run(fixtures.multiStep); expect(overBudget(run.segments, run.window)).toEqual([]);});Enforce in CI, alert in production. Failing a user request because a schema grew is worse than the degradation you're guarding against.
The discipline
Prompt engineering asks how to say it. Context engineering asks what deserves to be said at all, given a fixed budget and a model that reads unevenly.
Working on something like this?
Start a Conversation