CCAF logo

Domain 4 · Task 4.1

System Prompts with Explicit Criteria

Design prompts with explicit criteria to improve precision and reduce false positives.

When a reviewer or classifier produces too many false positives, the instinct is to tell it to be more cautious. That never works. The reliable fix is to replace vague instructions with explicit categorical criteria: state exactly what to flag, exactly what to skip, and give concrete code examples per severity level. This is the single most-tested idea in Domain 4 and it hinges on one discriminator — vague confidence language does not improve precision; explicit criteria do.

Key concept

A precision problem is a criteria problem. Replace vague language ("be conservative," "only high-confidence") with explicit categorical report-vs-skip rules plus concrete examples. "Be more conservative" is never the answer.

What you need to know

Explicit criteria beat vague instructions

Compare two ways of asking a reviewer to check comments:

  • Vague: "check that comments are accurate," "be conservative," "only report high-confidence findings."
  • Explicit: "flag a comment only when its claimed behavior contradicts the actual code behavior; do not flag stylistically outdated or merely incomplete comments."

The vague version gives the model no boundary to reason against, so it flags anything vaguely suspect. The explicit version defines the exact condition for a finding, which is what actually raises precision. General confidence language ("be conservative," "only high-confidence") does not improve precision — it only changes tone, not the decision boundary.

text
REPORT a finding only when ALL of these hold:
- It is a bug (incorrect behavior) or a security issue (injection, auth bypass, leaked secret).
- The claimed behavior contradicts the actual code behavior.

SKIP (do NOT report):
- Stylistic or formatting preferences.
- Locally-established patterns that are consistent within the file.
- Comments that are merely outdated in tone but not contradicted by the code.

Severity levels with concrete examples

Consistent classification needs explicit severity criteria with concrete code examples per level, not adjectives. Defining "critical vs high vs medium vs low" in the abstract produces drift; anchoring each level to a worked example produces repeatable judgments.

SeverityExplicit criterionConcrete example
criticalExploitable security flaw or data lossSQL built by string concatenation from user input
highIncorrect result in a common pathOff-by-one that drops the last record
mediumIncorrect result in an edge caseUnhandled empty-list input
lowCorrect but fragileMissing timeout on a network call

Restore trust by disabling noisy categories

A high false-positive category undermines trust in the accurate categories — once developers learn that "comment accuracy" findings are usually wrong, they start ignoring the security findings too. The tactical move is to temporarily disable the high-false-positive category while you rewrite its criteria, keeping the trustworthy categories live. This preserves signal on the categories that work instead of letting one bad category poison the whole reviewer. It is a scoping decision, not a confidence-threshold tweak.

Exam traps

The trapThe reality
The reviewer flags too many false positives, so tell it to "be more conservative" or "only report high-confidence findings."Vague confidence language does not improve precision. Define explicit categorical criteria for what to report vs skip, with concrete examples.
Raising the confidence threshold (e.g. only report findings above 0.9) filters out the false positives.Confidence-based filtering suppresses real findings alongside noise and leaves the decision boundary undefined. Categorical report-vs-skip criteria are the fix.
One noisy category is harmless as long as the other categories are accurate.A high false-positive category undermines trust in the accurate ones. Temporarily disable it while you improve its prompt.
Severity levels can be defined with adjectives like "serious" or "minor" and the model will classify consistently.Abstract adjectives cause drift. Anchor each severity level to concrete code examples for repeatable classification.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your CI pipeline runs Claude Code to review pull requests, but the team reports too many false positives — Claude flags minor style preferences and local patterns as if they were bugs, eroding developer trust in the review output. The current system prompt says: "Review the code thoroughly and be conservative about flagging issues, only reporting high- confidence findings."

What's the most effective change?

AHave Claude self-report a confidence score per finding and post-filter to only show findings above 0.85 confidence.
BRewrite the prompt with specific categorical criteria: define which categories of issues to report (bugs, security vulnerabilities, data flow errors) versus skip (minor style, local naming patterns, opinions).Correct
CStrengthen the existing instruction with stronger language: "Be EXTREMELY conservative. Only report issues you are 99% certain are real bugs."
DAdd more context about the codebase so Claude better understands which patterns are acceptable in this project.

Why: Task Statement 4.1 is explicit: general instructions like "be conservative" or "only report high-confidence findings" fail to improve precision compared to specific categorical criteria. The fix is to define which categories to report (bugs, security) versus skip (minor style, local patterns). Option A (confidence threshold filtering) doesn't help because the model's confidence is already miscalibrated. Option C (stronger emphatic language) is a degree of the same flawed approach. Option D (more context) helps but is downstream of having unclear criteria.

A customer-support agent has been in production for 4 months at 78% first-contact resolution. After your team pushed an update to the system prompt last week adding "be more conservative about issuing refunds," telemetry shows resolution dropped to 61%. The drop is concentrated in cases where the customer has clear policy entitlement to a refund — the agent now escalates these to humans even though it has clear authorization. Examining the system prompt, the new conservative-refund language was added without removing the existing "resolve customer issues efficiently" instruction. Both instructions now compete in every refund case.

ARevert the prompt change and instead add few-shot examples showing the agent processing routine refunds while escalating ambiguous cases.
BStrengthen the conservative-refund language with stronger emphatic words ("CRITICAL: only issue refunds with absolute certainty") so the model prioritizes it over the efficiency instruction.
CImplement a PreToolUse hook on process_refund that requires policy-check verification before execution; remove both competing instructions from the system prompt.
DAdd explicit categorical criteria to the prompt distinguishing situations where the agent should resolve versus escalate, with reasoning text showing which instruction takes precedence in each.Correct

Why: The competing-instructions problem. Reverting (A) loses the legitimate intent of the conservative-refund change. Emphatic words (B) don't resolve the actual contradiction — both instructions remain in conflict. A PreToolUse hook (C) treats this like a deterministic compliance problem, but the question is actually about decisional clarity, not enforcement. Categorical criteria with reasoning text (D) addresses the root cause: tell the model when to apply each instruction.

Build exercise

Turn a noisy code reviewer into a precise one

~45 min
  1. 1
    Take a review prompt that says "check the code carefully and report any problems, be conservative" and run it against a PR with a mix of real bugs and stylistic quirks.

    You should see: A high false-positive rate — style nits and locally-consistent patterns get flagged alongside the real bug.

  2. 2
    Rewrite the prompt with an explicit REPORT list (bugs, security) and an explicit SKIP list (style, local patterns, outdated-but-not-contradicted comments).

    Why: Explicit categorical criteria define the decision boundary; vague confidence language does not.

  3. 3
    Add a severity rubric with one concrete code example per level (critical/high/medium/low).

    Why: Anchoring each level to an example removes classification drift.

  4. 4
    Identify the single category producing the most false positives and temporarily disable it, leaving the trustworthy categories active.

    Why: One noisy category erodes trust in the accurate ones; disabling it preserves signal while you iterate.

  5. 5
    Re-run against the same PR and compare precision before and after.

    You should see: Real bugs still surface, but style nits and locally-consistent patterns no longer do.

Sources

Drill Prompt Engineering & Structured Output

Practice only this domain’s questions, untimed, with instant explanations.