CCAF logo

Domain 4 · Task 4.2

Few-Shot Prompting

Apply few-shot prompting to improve consistency and quality.

Few-shot prompting — including a handful of input/output examples — is the most effective technique for getting consistently formatted, actionable output when prose instructions produce inconsistent results. An example unambiguously demonstrates both the output format and the decision logic, and the model generalizes the pattern to novel cases rather than merely echoing the examples you gave. Use it for inconsistent format, ambiguous judgment, and novel patterns — but not for a problem that actually needs a hard gate.

Key concept

Inconsistent format, ambiguous judgment, or novel patterns → few-shot (2–4 examples that show the reasoning). For a mandatory-ordering problem, few-shot is the wrong tool — use a gate.

What you need to know

Why examples beat prose

Prose describes the target; an example shows it. Few-shot is more effective than instructions alone because a single example unambiguously fixes both the format (what fields, in what shape) and the decision logic (which action to take on a borderline case). Three concrete payoffs:

  • Ambiguous-case handling — the examples demonstrate what to do at the edges, where instructions are vague.
  • Generalization to novel patterns — the model learns the underlying rule and applies it to cases you never showed, rather than pattern-matching the literal examples.
  • Reduced hallucination in extraction — showing how to handle informal measurements and varied structures curbs invented values and empty/null required fields.

Design 2–4 examples that show the reasoning

Create 2–4 targeted examples — not a dozen. Aim them at ambiguous scenarios and make each one show the reasoning for choosing one action over a plausible alternative, not just the final answer. Cover:

  • The desired output format (location, issue, severity, suggested fix).
  • Acceptable patterns vs genuine issues, so the model reduces false positives while still generalizing.
  • Varied document structures — e.g. inline citations vs a bibliography — and the fix for empty/null extraction of required fields.

Pair the examples with format-normalization rules in the prompt (dates → ISO 8601, currency → amount + code) to curb semantic errors.

text
Example 1 (ambiguous tool selection — show the reasoning):
Request: "Can you clean this up?" on a config file.
Reasoning: "clean up" is ambiguous — it could mean format or refactor.
The file is a config, so formatting is the safe, reversible action.
Action: call format_file (NOT refactor_code).

Example 2 (informal measurement — normalize, don't drop):
Text: "about a cup and a half of flour"
Extraction: { "ingredient": "flour", "amount": 1.5, "unit": "cup", "approximate": true }

When few-shot is the wrong tool

Few-shot fixes judgment and format problems. It does not enforce a hard constraint. If the requirement is a *mandatory ordering* — "validation must always run before the write" — examples only make the model *usually* comply, which is not the same as *always*. That is a gate (a deterministic check or hook in the control flow), not a prompting problem. On the exam, a scenario about inconsistent output format or inconsistent tool selection points to few-shot; a scenario about a step that must never be skipped points to a gate.

Exam traps

The trapThe reality
The output format keeps varying, so add more detailed prose describing the exact format.Prose alone is what produced the inconsistency. 2–4 examples showing the format are the most effective fix.
Few-shot examples just teach the model to match the exact cases you provided.Examples enable generalization to novel patterns — the model learns the underlying rule, not just the literal cases.
More examples are always better, so include 15–20 to cover every case.2–4 targeted examples that show the reasoning are the sweet spot; more adds cost and dilutes focus without proportionate gain.
Few-shot can guarantee a mandatory step (like validation before write) always runs.Few-shot shapes judgment, not hard constraints. A mandatory ordering needs a gate, not examples.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

You're processing 30K research papers per night for structured extraction. Schema validation is at 100% (tool_use with strict schemas). But analysis of the outputs shows ~7% of extracted methodology_section fields contain summaries of the abstract rather than the methodology section. The error happens specifically in papers where the methodology details are embedded within the introduction rather than in a separately-labeled section. The prompt currently says "extract methodology details from the methodology section."

AAdd a verification pass with a second model checking each extraction against the source.
BAdd few-shot examples demonstrating extractions from papers with methodology embedded in introductions, showing how to recognize the pattern and extract correctly.Correct
CMake methodology_section optional in the schema and flag papers without a labeled methodology section for human review.
DSwitch to a more capable model tier with better long-document understanding.

Why: Few-shot examples that demonstrate the specific failure mode (methodology embedded in introductions) train the model to recognize the pattern. Verification passes (A) are expensive. Optional fields + human review (C) papers over the symptom. Model upgrade (D) doesn't teach the recognition pattern.

Your schema includes a skills: string[] field. Production monitoring reveals three consistency issues: (1) compound phrases like "Python and SQL" are sometimes kept as one entry, sometimes split; (2) implied but unstated skills occasionally appear in extractions; (3) similar documents produce wildly different array lengths (5-10 vs 40+ entries). Your prompt currently says "Extract all skills mentioned." What's the most effective improvement?

AAdd few-shot examples demonstrating compound phrase handling, explicit mention criteria, and appropriate entry granularity.Correct
BAdd constraints: "Extract 10-20 skills maximum, one skill per entry, only explicitly named skills."
CAdd post-extraction normalization that maps skills to a canonical taxonomy and deduplicates similar entries.
DEnrich the schema to {skill: string, confidence: float, source_quote: string}[] to capture extraction metadata.

Why: All three issues are about the model's interpretation of what counts as 'a skill.' Few-shot examples teach the pattern concretely — split vs. not-split, mentioned vs. inferred, appropriate granularity.

Build exercise

Stabilize an inconsistent extractor with few-shot examples

~40 min
  1. 1
    Prompt Claude to extract issues from code review comments using only prose instructions, and run it across several comments with varied phrasing.

    You should see: The output shape drifts — sometimes a severity, sometimes not; informal cases get dropped or invented.

  2. 2
    Add 2–4 examples covering an ambiguous case, an informal-measurement case, and a varied-structure case — each showing the reasoning, not just the answer.

    Why: Examples that show reasoning teach the decision logic so the model generalizes to novel cases.

  3. 3
    Add format-normalization rules (dates → ISO 8601, currency → amount + code) alongside the examples.

    Why: Normalization rules curb the semantic errors that examples alone don't fully catch.

  4. 4
    Include one example that returns null for a genuinely-absent required field.

    Why: Demonstrating the null case fixes empty/fabricated extraction of required fields.

  5. 5
    Re-run the full set and confirm the output shape is now consistent, then test a phrasing you never showed.

    You should see: Consistent format across all inputs, including the novel phrasing — evidence of generalization.

Sources

Drill Prompt Engineering & Structured Output

Practice only this domain’s questions, untimed, with instant explanations.