Domain 4 · Task 4.6
Multi-Instance & Multi-Pass Review
Design multi-instance and multi-pass review architectures.
Two architectural moves define reliable review. First, a model that generated code keeps its generation reasoning and is less likely to question its own decisions in the same session — so an independent review instance with no prior context catches subtle issues that self-review and extended thinking miss. Second, a single giant review over many files suffers attention dilution and produces contradictory findings — so you split it into per-file local passes plus a cross-file integration pass. Neither problem is solved by a bigger model or a bigger context window.
Key concept
Reviewing generated code → an independent instance, not the same session. Many files → per-file passes + a separate integration pass, not one big pass and not a bigger model.
What you need to know
Why self-review is weak
When the same session that generated code then reviews it, it carries the reasoning that produced the code — including the assumptions that led to any bug. Sharing that context makes the model less likely to question its own decisions: it already "knows" why the code is the way it is. This is why extended thinking in the generating session does not fix it either — more thinking over the same biased context still defends the same decisions. A second, independent Claude instance reviews the code cold, with no generation reasoning to anchor on, and catches the subtle issues the generator rationalized away.
code = generate_session.run("Implement the payment handler.")
# Independent instance — NO generation context, reviews cold.
review_session = client.new_session()
findings = review_session.run(
"Review this code for bugs and security issues. "
"You did not write it; do not assume it is correct.\n\n" + code
)Per-file passes plus an integration pass
A large multi-file review — say a 14-file PR — done as one big pass dilutes the model's attention across too much context, producing inconsistent and contradictory findings. The fix is prompt chaining the review into focused passes:
- Per-file local passes — review each file on its own, so full attention lands on that file's logic.
- A separate integration pass — review the cross-file concerns (interfaces, contracts, data flow between the files) that a per-file pass can't see.
This splits a diluted single judgment into several focused ones plus one that specifically owns the seams between files.
Confidence-scored verification passes
To route reviewer attention efficiently, run a verification pass where the model self-reports its confidence per finding. Low-confidence findings get human attention or a second look; high-confidence findings can flow through. This is a routing mechanism, not a suppression filter — critically, it is different from "only report a finding if it appears in ≥2 of 3 independent runs," which suppresses real bugs that happen to surface in just one run. Confidence self-reporting routes attention; consensus voting silently drops genuine findings.
Exam traps
| The trap | The reality |
|---|---|
| The session that wrote the code can review it well if you turn on extended thinking. | The generating session keeps its reasoning and won't question its own decisions. Use a second independent instance with no generation context. |
| A 14-file PR review is inconsistent because the context window is too small — use a bigger context window. | A bigger window doesn't fix attention quality. Split into per-file passes plus a separate integration pass. |
| A bigger or smarter model will fix contradictory multi-file review findings. | The problem is attention dilution across too much context, not model capability. Prompt-chain into focused per-file and integration passes. |
| Only reporting findings that appear in at least 2 of 3 runs improves review quality. | Consensus voting suppresses real bugs that surface in a single run. Use per-finding confidence self-reporting to route attention instead. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
Your CI runs Claude Code reviews on PRs. Reviews must complete within 8 minutes (CI budget). The team's PRs average 12 files but range up to 60 files. For PRs above ~20 files, you're seeing 2 issues: (1) reviews sometimes exceed 8 minutes and CI fails them, (2) review quality drops — Claude misses cross-file bugs and produces contradictory feedback (flagging a pattern as bad in one file but approving it in another in the same PR).
Why: Multi-pass review (Task 4.6): per-file local passes + cross-file integration pass. Large-PR attention dilution is the root cause; bigger context windows don't fix attention quality. Sampling (C) misses real bugs. Temperature (D) is orthogonal.
Production metrics show that when your agent resolves complex cases involving billing disputes or multi-order returns, customer satisfaction scores are 15% lower than for simple cases—even when the resolution is technically correct. Root cause analysis reveals the agent provides accurate resolutions but inconsistently explains the reasoning: sometimes omitting relevant policy details, other times missing timeline information or next steps. The specific context gaps vary by case. You want to improve resolution quality without adding human review overhead. Which approach is most effective?
Why: Self-critique (evaluator-optimizer pattern) directly addresses the root cause: inconsistent inclusion of explanation elements. By having the agent evaluate its own response against specific criteria before presenting it, you catch case-specific gaps that vary with each situation. Few-shot examples (B) help with consistent patterns but cannot cover highly variable gaps. Model tier upgrades (C) don't address the structural issue—the resolutions are already accurate, just incompletely explained. Confirmation steps (D) shift the burden to customers rather than improving the agent's output quality.
Build exercise
Design an independent, multi-pass review pipeline
~50 min- 1Generate a small multi-file feature in one Claude session, then ask that same session to review its own code.
You should see: The self-review misses or rationalizes subtle issues it introduced.
- 2Start a fresh, independent instance with no generation context and have it review the same code cold.
Why: An instance without the generator's reasoning is more willing to question the decisions.
- 3For a larger PR, run a separate per-file review pass on each file individually.
Why: Per-file passes give full attention to each file's logic without dilution.
- 4Add a dedicated integration pass that reviews only the cross-file interfaces and data flow.
Why: Cross-file concerns are invisible to per-file passes and need their own pass.
- 5Add a verification pass where the model self-reports confidence per finding, and route low-confidence findings to a human.
Why: Confidence self-reporting routes attention without suppressing single-run real bugs the way consensus voting does.
- 6Compare the independent multi-pass results against the original single self-review pass.
You should see: More real issues found and fewer contradictions than the one-big-pass self-review.
Sources
Drill Prompt Engineering & Structured Output
Practice only this domain’s questions, untimed, with instant explanations.