CCAF logo

Domain 5 · Task 5.4

Codebase Exploration & Context Degradation

Manage context effectively in large-codebase exploration.

Exploring a large codebase is a context-hungry task: dozens of files, each read appended to history, and after a while the window is full of raw discovery. The symptom of context degradation is unmistakable — the agent gives inconsistent answers and starts citing "typical patterns" instead of the specific classes it found earlier. This lesson covers the four tools that fix it: subagent delegation, scratchpad files, pre-spawn summarization, and structured state for crash recovery.

Key concept

Long exploration → offload to subagents plus scratchpad plus summaries; don't let raw discovery fill the main window.

What you need to know

Recognizing context degradation

The tell-tale sign, and the exam's framing, is: after exploring 30+ files the agent gives inconsistent answers and cites generic "typical patterns" rather than the specific classes and functions it actually found. This is the main context window filling with verbose file contents until earlier findings are lost or diluted. The wrong response is "just keep going" or "use a bigger context window" — both leave the raw discovery in the window. The right response restructures where the exploration output lives.

Subagent delegation for verbose discovery

Spawn subagents for specific questions — "find all test files", "trace the refund-flow dependencies" — while the main agent keeps only high-level coordination. Because a subagent runs in an isolated context, its verbose file reading never touches the coordinator's window; only the distilled answer comes back. This is the core move: the expensive, noisy work happens in a context you can discard, and the main agent stays focused on the plan.

Scratchpad files and pre-spawn summaries

Persist key findings in scratchpad files on disk and reference them as you go, so a finding from file 3 survives long after file 30. Before spawning the next phase's subagents, summarize the key findings and inject that summary into their initial context — so each phase starts with the distilled state rather than re-deriving it. When the main window fills with verbose discovery despite this, use /compact to condense it. The through-line: important findings belong in durable, referenceable form, not only in the volatile conversation.

Structured state persistence for crash recovery

For long, multi-phase explorations, design crash recovery through structured state exports: each agent exports its state, and on resume the coordinator loads a manifest describing what has been done and found. This makes the exploration resumable across context boundaries and process crashes, rather than starting over. It is the same principle as the scratchpad, formalized: state lives outside any single context window so no boundary — compaction, crash, or new session — loses it.

Exam traps

The trapThe reality
The agent is giving inconsistent answers after exploring many files, so give it a bigger context window or a bigger model.More capacity still fills with raw discovery and still degrades. Offload verbose exploration to subagents and persist findings in scratchpad files.
Just keep exploring in the same session — the model will remember what it found earlier.That's exactly the degradation scenario. Summarize before the next phase, use scratchpad files, and delegate to subagents to keep the main window clean.
Subagents share the coordinator's context, so they'll already know what's been found.Subagents run in isolated contexts and inherit nothing but the prompt you pass. Inject the summarized findings into their initial context explicitly.
If the session crashes mid-exploration, there's no way to resume without redoing the work.Structured state exports plus a manifest the coordinator loads on resume make exploration crash-recoverable.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

A platform team is debugging an issue: their Claude Code agent, working in a 60-file refactor, occasionally references file content that no longer exists — for example, it suggests editing a function that was renamed three turns ago. Investigation reveals:

The session is long (40+ turns)

The agent re-reads files only when explicitly prompted to

Auto-compaction kicked in at turn 35 and the team can see in the transcript that compaction summarized away some of the file-read history

The renamed function was renamed at turn 20, re-read at turn 22, then compacted at turn 35

Two senior engineers propose different fixes. Which is most likely to address the root cause without trading one problem for another?

AIncrease the auto-compaction threshold so compaction happens less often — the agent will have the file contents in context for longer
BDisable auto-compaction entirely; rely on the user to manually /clear and provide a summary when the session is getting long
CAdd a PreCompact hook that writes the current file-read state to a scratchpad file before compaction, and add a SessionStart /post-compact behavior where the agent re-reads files it had recently been working with — so post-compaction state stays accurateCorrect
DMove to a model with a larger context window so compaction never has to happen

Why: This is a subtle reliability question. The symptom (stale state after compaction) isn't fixed by: A (delay compaction): Eventually you hit the threshold anyway; the problem returns. B (disable compaction): Trades automatic management for manual /clear — and now the user has to do summarization manually, which they'll often do worse than the system. D (bigger window): Pushes the problem further out but the same degradation will appear at the new ceiling. C is the engineered solution: use the PreCompact hook to persist the relevant working state to disk before compaction, then have post-compaction behavior re-establish the agent's working knowledge from the persisted state. This is the documented pattern for sessions that need durable working memory across compaction events.

Your developer productivity agent has been exploring a large legacy codebase for 35 turns. It read 60+ files and built detailed understanding of several subsystems. By turn 40, you observe that it references "typical patterns" rather than the specific class names it discovered earlier (e.g., it says "the standard repository pattern" instead of citing the PaymentRepository class it analyzed at turn 12).

What's the most effective approach?

ASwitch to a higher-tier model with a larger context window.
BHave the agent maintain a scratchpad file that records key findings (class names, file locations, relationships) and reference it explicitly when asked about earlier discoveries.Correct
CUse /compact to reduce context usage so the model has more attention budget for remaining exploration.
DReset the session with a fresh start, manually providing a summary of the prior exploration's key findings.

Why: Task Statement 5.4 specifically describes this: have agents maintain scratchpad files recording key findings and reference them for subsequent questions to counteract context degradation in extended sessions. The symptom — referencing "typical patterns" rather than specific classes — is exactly the documented signal that prompts scratchpad usage. Option A (larger model) is expensive and doesn't fix the underlying attention degradation pattern. Option C ( /compact ) helps with context length but discards the specific findings you want to preserve. Option D (reset + summary) loses fidelity that the scratchpad approach keeps verbatim.

Build exercise

Explore a large repo without degrading context

~50 min
  1. 1
    Pick a repo of 30+ source files and, in a single session, ask the agent to read them one by one and then answer a specific question about a class it saw early on.

    Why: You need enough files that the main window fills and earlier findings start to blur.

  2. 2
    Record the baseline: note where the agent starts giving inconsistent answers or citing "typical patterns" instead of the actual class.

    You should see: Degradation appears — generic answers replacing specific, file-grounded ones.

  3. 3
    Restructure so a subagent handles each targeted question ("find all test files", "trace the refund flow") and returns only a distilled answer to the main agent.

    Why: Isolating verbose reading in subagent contexts keeps the coordinator window clean.

  4. 4
    Have each phase write its key findings to a scratchpad file and reference those files on later questions.

    You should see: Findings from early files remain accurate late in the session because they're read back from disk.

  5. 5
    Before spawning the next exploration phase, summarize the findings so far and inject that summary into the subagents' initial context.

    Why: Each phase starts from distilled state instead of re-deriving it from raw files.

  6. 6
    Export a manifest of completed work and findings, kill the session, and resume from the manifest.

    You should see: The resumed session picks up from the recorded state rather than starting the exploration over.

Sources

Drill Context Management & Reliability

Practice only this domain’s questions, untimed, with instant explanations.