CCAF logo

Domain 5 · Task 5.1

Context Window Management

Manage conversation context to preserve critical information.

The context window holds everything the model can see at once: system prompt, full message history, tool definitions, and tool results. In a long conversation that budget is under constant pressure, and the naive fixes — summarizing history, letting tool outputs pile up — quietly corrupt exactly the details that matter. This lesson is about preserving critical information across a long interaction by controlling where it lives, not by hoping it survives compression.

Key concept

Hard facts live in a persistent structured block, not in prose history that gets summarized. Trim tool outputs at the source.

What you need to know

Three failure modes of a long context

Three things go wrong as a conversation grows, and the exam expects you to name them:

  • Progressive-summarization drift. Compressing history to save tokens turns exact numbers, percentages, and dates into vague language — "$89.99" becomes "about ninety dollars", "2025-01-15" becomes "mid-January". The customer's stated expectations blur the same way.
  • Lost-in-the-middle. Models use information at the start and end of a long input reliably but may omit material buried in the middle. Position, not just presence, decides whether the model attends to a fact.
  • Tool-result accumulation. Every tool call appends its full output to history. A tool that returns 40+ fields when only 5 matter burns the window and adds noise that crowds out signal.

The case-facts block

The fix for summarization drift is to stop trusting prose to carry the numbers. Extract transactional facts — amounts, dates, order numbers, statuses, the customer's explicit request — into a persistent case-facts block that is included verbatim in every prompt, positioned outside the summarized history. Because it is re-sent unchanged each turn, it never degrades. For multi-issue sessions, persist structured issue data into a separate context layer rather than letting several problems tangle in one narrative.

text
=== CASE FACTS (always included, outside summarized history) ===
Customer ID: CUST-12345 | Order: ORD-67890 | Date: 2025-01-15
Amount: $89.99 | Issue: damaged on delivery | Request: full refund
Status: pending manager approval

Trim tool outputs at the source

Accumulation is best solved where the output is produced, not after the window is already full. Trim verbose tool results to the relevant fields before they enter history — often with a PostToolUse hook that strips a 40-field payload down to the 5 fields the workflow uses. Complementary tactics: place key-findings summaries at the beginning and use explicit section headers to counter lost-in-the-middle; require subagents to include metadata (dates, sources, methodology) in structured outputs; and when downstream budgets are tight, have upstream agents return structured data (facts, citations, relevance scores) instead of verbose reasoning.

Exam traps

The trapThe reality
Progressive summarization is fine for keeping the conversation coherent, so critical numbers and dates will survive it.Summarization is exactly what degrades exact values into vague phrases. Keep hard facts in a persistent case-facts block that is never summarized.
Using a bigger context window (or a bigger model) solves the forgetting problem.A larger window still suffers lost-in-the-middle and still fills with accumulated tool noise. Restructure where information lives; don't just add capacity.
Putting the important case details somewhere in the conversation is enough — the model will find them.Material in the middle of a long input is often omitted. Position key findings at the start/end and use explicit section headers.
Every field a tool returns is worth keeping in case it's needed later.Unfiltered 40-field outputs consume the window out of proportion to relevance. Trim to the fields that matter, ideally at the source via a PostToolUse hook.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your customer support agent has a 12K-token system prompt (policies, tool descriptions, few- shot examples). The agent handles ~50K conversations daily, each averaging 5 turns. You enable prompt caching with cache_control on the system prompt.

Which statement is most accurate about how this caching behaves?

ACached prompts persist indefinitely on Anthropic's servers; once cached, the system prompt never needs to be re-sent
BCache hits are scoped to the same conversation only — across separate conversations, each pays the full system-prompt cost on its first turn
CCache hits occur across any requests with matching prefixes; cached reads cost a fraction of base input; entries expire after ~5 minutes of inactivity by default (extended TTL available at higher write cost)Correct
DCaching is automatic for any prompt over 4K tokens; cache_control only controls cache eviction priority, not whether caching happens

Why: Prompt caching matches prefixes across requests — any conversation reusing the same system prompt benefits, not just turns within one conversation (B is wrong). Cached reads cost a fraction of base input; entries expire after ~5 minutes of inactivity by default (an extended TTL is available at a higher write cost). Caching is not permanent (A) and not automatic for large prompts (D) — the cache_control marker tells the API where to set the cache breakpoint, which is what activates caching at that boundary.

Your system prompt has four logical sections of substantial size: (1) overall instructions (2KB), (2) tool definitions (5KB), (3) policy documents (8KB), (4) few-shot examples that vary per use case (4KB). Sections 1–3 are stable across requests; section 4 changes frequently.

Where do you place cache_control breakpoints to maximize cache hits?

AOne breakpoint at the very end of section 4 — caching everything together gives the largest single cached block
BUp to 4 breakpoints are allowed; placing one at the end of section 3 caches the stable 15KB while allowing section 4 to vary per request without invalidating the cacheCorrect
CPlace a breakpoint at the start of each of the four sections to create independent cache entries
DCaching only works at the boundary between system prompt and user messages — intermediate breakpoints are invalid

Why: Each cache_control breakpoint caches everything up to and including that point. The API supports up to 4 breakpoints per request. Placing one at the end of section 3 caches the stable 15KB; section 4 can vary per request without invalidating the cache. Caching everything (A) means every change to section 4 invalidates the entire cache. Per-section breakpoints (C) waste the 4-breakpoint budget — each entry must still match its full prefix to hit, so the smaller cached entries don't help when used in different contexts. D is false — intermediate breakpoints are explicitly supported.

Build exercise

Preserve case facts across a long support conversation

~45 min
  1. 1
    Simulate a support session: seed a conversation with a customer ID, an order number, a refund amount, a date, and an explicit request, then pad it with 20+ turns of back-and-forth.

    Why: You need enough length that naive summarization would kick in and blur the exact values.

  2. 2
    First run a baseline that summarizes older turns into prose, then ask the agent to restate the exact refund amount and order date.

    You should see: The summarized run answers with vague or drifted values ("around $90", "mid-January") — the drift failure made visible.

  3. 3
    Add a case-facts block containing the transactional facts and inject it verbatim into every request, outside the summarized history.

    Why: Re-sending an unchanged structured block is what makes the facts immune to summarization.

  4. 4
    Re-run the same restate question with the case-facts block present.

    You should see: The agent returns the exact amount, order number, and date every time, regardless of conversation length.

  5. 5
    Wire a PostToolUse step that trims a verbose 40-field tool response down to the 5 relevant fields before it enters history.

    Why: Trimming at the source stops tool-result accumulation from crowding out the signal.

  6. 6
    Move the case-facts block to the very start of the prompt and add a section header, then confirm retrieval stays reliable even in the longest transcripts.

    Why: Start/end placement plus headers mitigates lost-in-the-middle on top of the persistence guarantee.

Sources

Drill Context Management & Reliability

Practice only this domain’s questions, untimed, with instant explanations.