Domain 5 · Task 5.1
Context Window Management
Manage conversation context to preserve critical information.
The context window holds everything the model can see at once: system prompt, full message history, tool definitions, and tool results. In a long conversation that budget is under constant pressure, and the naive fixes — summarizing history, letting tool outputs pile up — quietly corrupt exactly the details that matter. This lesson is about preserving critical information across a long interaction by controlling where it lives, not by hoping it survives compression.
Key concept
Hard facts live in a persistent structured block, not in prose history that gets summarized. Trim tool outputs at the source.
What you need to know
Three failure modes of a long context
Three things go wrong as a conversation grows, and the exam expects you to name them:
- Progressive-summarization drift. Compressing history to save tokens turns exact numbers, percentages, and dates into vague language — "$89.99" becomes "about ninety dollars", "2025-01-15" becomes "mid-January". The customer's stated expectations blur the same way.
- Lost-in-the-middle. Models use information at the start and end of a long input reliably but may omit material buried in the middle. Position, not just presence, decides whether the model attends to a fact.
- Tool-result accumulation. Every tool call appends its full output to history. A tool that returns 40+ fields when only 5 matter burns the window and adds noise that crowds out signal.
The case-facts block
The fix for summarization drift is to stop trusting prose to carry the numbers. Extract transactional facts — amounts, dates, order numbers, statuses, the customer's explicit request — into a persistent case-facts block that is included verbatim in every prompt, positioned outside the summarized history. Because it is re-sent unchanged each turn, it never degrades. For multi-issue sessions, persist structured issue data into a separate context layer rather than letting several problems tangle in one narrative.
=== CASE FACTS (always included, outside summarized history) ===
Customer ID: CUST-12345 | Order: ORD-67890 | Date: 2025-01-15
Amount: $89.99 | Issue: damaged on delivery | Request: full refund
Status: pending manager approvalTrim tool outputs at the source
Accumulation is best solved where the output is produced, not after the window is already full. Trim verbose tool results to the relevant fields before they enter history — often with a PostToolUse hook that strips a 40-field payload down to the 5 fields the workflow uses. Complementary tactics: place key-findings summaries at the beginning and use explicit section headers to counter lost-in-the-middle; require subagents to include metadata (dates, sources, methodology) in structured outputs; and when downstream budgets are tight, have upstream agents return structured data (facts, citations, relevance scores) instead of verbose reasoning.
Exam traps
| The trap | The reality |
|---|---|
| Progressive summarization is fine for keeping the conversation coherent, so critical numbers and dates will survive it. | Summarization is exactly what degrades exact values into vague phrases. Keep hard facts in a persistent case-facts block that is never summarized. |
| Using a bigger context window (or a bigger model) solves the forgetting problem. | A larger window still suffers lost-in-the-middle and still fills with accumulated tool noise. Restructure where information lives; don't just add capacity. |
| Putting the important case details somewhere in the conversation is enough — the model will find them. | Material in the middle of a long input is often omitted. Position key findings at the start/end and use explicit section headers. |
| Every field a tool returns is worth keeping in case it's needed later. | Unfiltered 40-field outputs consume the window out of proportion to relevance. Trim to the fields that matter, ideally at the source via a PostToolUse hook. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
Your customer support agent has a 12K-token system prompt (policies, tool descriptions, few- shot examples). The agent handles ~50K conversations daily, each averaging 5 turns. You enable prompt caching with cache_control on the system prompt.
Which statement is most accurate about how this caching behaves?
cache_control only controls cache eviction priority, not whether caching happensWhy: Prompt caching matches prefixes across requests — any conversation reusing the same system prompt benefits, not just turns within one conversation (B is wrong). Cached reads cost a fraction of base input; entries expire after ~5 minutes of inactivity by default (an extended TTL is available at a higher write cost). Caching is not permanent (A) and not automatic for large prompts (D) — the cache_control marker tells the API where to set the cache breakpoint, which is what activates caching at that boundary.
Your system prompt has four logical sections of substantial size: (1) overall instructions (2KB), (2) tool definitions (5KB), (3) policy documents (8KB), (4) few-shot examples that vary per use case (4KB). Sections 1–3 are stable across requests; section 4 changes frequently.
Where do you place cache_control breakpoints to maximize cache hits?
Why: Each cache_control breakpoint caches everything up to and including that point. The API supports up to 4 breakpoints per request. Placing one at the end of section 3 caches the stable 15KB; section 4 can vary per request without invalidating the cache. Caching everything (A) means every change to section 4 invalidates the entire cache. Per-section breakpoints (C) waste the 4-breakpoint budget — each entry must still match its full prefix to hit, so the smaller cached entries don't help when used in different contexts. D is false — intermediate breakpoints are explicitly supported.
Build exercise
Preserve case facts across a long support conversation
~45 min- 1Simulate a support session: seed a conversation with a customer ID, an order number, a refund amount, a date, and an explicit request, then pad it with 20+ turns of back-and-forth.
Why: You need enough length that naive summarization would kick in and blur the exact values.
- 2First run a baseline that summarizes older turns into prose, then ask the agent to restate the exact refund amount and order date.
You should see: The summarized run answers with vague or drifted values ("around $90", "mid-January") — the drift failure made visible.
- 3Add a case-facts block containing the transactional facts and inject it verbatim into every request, outside the summarized history.
Why: Re-sending an unchanged structured block is what makes the facts immune to summarization.
- 4Re-run the same restate question with the case-facts block present.
You should see: The agent returns the exact amount, order number, and date every time, regardless of conversation length.
- 5Wire a PostToolUse step that trims a verbose 40-field tool response down to the 5 relevant fields before it enters history.
Why: Trimming at the source stops tool-result accumulation from crowding out the signal.
- 6Move the case-facts block to the very start of the prompt and add a section header, then confirm retrieval stays reliable even in the longest transcripts.
Why: Start/end placement plus headers mitigates lost-in-the-middle on top of the persistence guarantee.
Sources
Drill Context Management & Reliability
Practice only this domain’s questions, untimed, with instant explanations.