CCAF logo

Domain 4 · Task 4.4

Validation, Retry & Feedback Loops

Implement validation, retry, and feedback loops for extraction quality.

tool_use gives you syntactically valid output; validation and retry give you semantically correct output. When an extraction fails a check — line items don't sum, a value is in the wrong field — the fix is retry with error feedback: send the model the original document, its failed extraction, and the specific error, and ask it to self-correct. The critical judgment is knowing the limit of retry: it fixes format, structural, and arithmetic errors, but it cannot recover information that is simply not present in the source.

Key concept

Format / structure / arithmetic error → retry with the specific error and the original document. Information absent from the source → retry won't help; get the missing source or route to a human.

What you need to know

Retry with error feedback

The retry loop re-prompts with three things: the original input + the failed output + the specific validation error. Naming the exact error is what makes self-correction work — "the line items sum to $145.00 but stated_total is $150.00" guides the correction far better than "try again." This is effective for semantic errors that survive the schema:

  • Values that don't add up (arithmetic).
  • A value placed in the wrong field (structural).
  • Format mismatches (a date that isn't ISO 8601).

Syntax errors don't need this loop — tool_use already eliminates them.

python
def extract_with_retry(doc, max_retries=2):
    result = extract(doc)
    for _ in range(max_retries):
        errors = validate(result)          # semantic checks: sums, fields, formats
        if not errors:
            return result
        if is_absent_from_source(errors):  # retry cannot conjure missing data
            return route_to_human(doc, errors)
        result = extract(
            doc,
            prior_output=result,
            feedback=f"Validation failed: {errors}. Re-extract, correcting these.",
        )
    return route_to_human(doc, errors)

The limit of retry: absent information

Retry is powerful but not omnipotent. If a required value is absent from the source — for instance the ship date lives in a separate document you never supplied — no amount of re-prompting will produce it, because the model has nothing to read. Distinguish the two cases:

  • Retry will help: the data is in the document but was mis-parsed, mis-placed, or mis-formatted.
  • Retry won't help: the data is genuinely not in the provided source. The fix is to supply the missing source or route to a human — not to loop.

A loop that keeps retrying on absent data either spins forever or, worse, pressures the model into fabricating the value.

Self-correction and feedback-loop design

You can make discrepancies surface deterministically by extracting a stated value alongside a computed one and a conflict flag, so validation is a simple boolean check instead of re-derivation. Add conflict_detected booleans for inconsistent sources. For feedback-loop design over time, add a detected_pattern field that records which construct triggered each finding — then, when developers dismiss findings, you can analyze the dismissal patterns and tighten the criteria for the categories that produce false positives.

json
{
  "stated_total": "$150.00",
  "calculated_total": "$145.00",
  "conflict_detected": true,
  "detected_pattern": "line_items_sum_mismatch"
}

Exam traps

The trapThe reality
Extraction failed validation because a field is missing — retry with feedback will recover it.If the field is absent from the provided source (e.g. it lives in an unsupplied document), retry can't conjure it. Get the source or route to a human.
Retrying with a generic "that was wrong, try again" is enough.Effective retry sends the original document, the failed output, and the specific error. Naming the exact error is what drives correction.
tool_use already validates the data, so a separate validation step is redundant.tool_use eliminates syntax errors only. Semantic checks (sums, field placement, formats) are a separate, necessary step.
To catch a total that doesn't match, re-derive it in downstream code.Have the model emit stated_total and calculated_total with a conflict_detected flag so the discrepancy surfaces deterministically at extraction time.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your structured-extraction pipeline uses tool_use with a strict JSON schema. After deployment, JSON syntax errors dropped to 0% — the schema enforces structure. But you see a new class of error: values that match the schema's type but are semantically wrong. Examples: a date field receives "2024-13-45" (string matches pattern but isn't a real date); a tax_rate field receives 0.85 (correct type, but 85% — implausibly high for any tax jurisdiction in your dataset).

What addresses this?

ASwitch from tool_use to JSON mode with a more permissive schema
BLower temperature to 0 to reduce semantic mistakes
CAdd programmatic validation on top of the structured output (real date parsing, range checks on numeric fields)Correct
DAdd more few-shot examples to the prompt showing correct semantic values

Why: tool_use with a JSON schema guarantees syntactic compliance (output will parse and field types match) but does not guarantee semantic correctness. "2024-13-45" matches a string field's type; 0.85 matches a number's type. Programmatic validation on top — real date parsing, plausibility checks on numeric ranges, business-rule validation — is the layer that catches semantic errors. Few-shots (D) help at the margin but aren't a guarantee. Temperature (B) doesn't change what the model considers semantically valid. Switching to JSON mode (A) is strictly worse than tool_use for reliability.

Your data extraction pipeline processes restaurant menus. The schema includes prices: number[] . After deployment, ~3% of extractions fail validation: prices come back as ["$12.99", "$15.50"] (strings with dollar signs) instead of [12.99, 15.50] (numbers). You implement a retry-with-error-feedback approach where on validation failure, you re-prompt the model with the source document, the failed extraction, and the specific error. Most failures resolve within 2-3 attempts. For which failure pattern would additional retries be LEAST likely to resolve the issue?

AModel returns ["$12.99", "$15.50"] when schema requires numbers.
BModel returns a 2D nested array when schema requires a flat array.
CModel returns prices in the wrong order: dinner prices before lunch prices, when the menu lists lunch first.
DModel returns [null, null, 12.99] because three menu items have prices listed only as "ask your server" in the source document.Correct

Why: Retries can't add information that isn't in the source. This is the documented Task 4.4 limit. A, B, C are all structural/format mismatches where the model has the right value but wrong shape — retry-with-feedback handles those.

Build exercise

Build a validation-and-retry extraction loop

~50 min
  1. 1
    Extract an invoice with tool_use, then write a validator that checks line items sum to the stated total and dates are ISO 8601.

    Why: Semantic validation catches the errors that tool_use lets through.

  2. 2
    On failure, re-prompt with the original document, the failed extraction, and the exact error string.

    You should see: The model self-corrects an arithmetic or format mistake on the next pass.

  3. 3
    Add a stated_total / calculated_total / conflict_detected pattern so mismatches surface as a boolean.

    Why: Deterministic conflict flags turn validation into a simple check instead of re-derivation.

  4. 4
    Feed a document whose required ship_date genuinely lives in a separate file you don't provide, and detect that retry doesn't resolve it.

    You should see: The value stays null across retries — the signal to stop looping and route elsewhere.

  5. 5
    Add a branch that routes the absent-data case to a human (or requests the missing source) instead of retrying.

    Why: Retry can't conjure missing information; the correct escalation is to get the source or a human.

  6. 6
    Cap retries (e.g. 2) so a persistently-failing extraction escalates rather than looping forever.

    Why: A cap prevents runaway loops and forces escalation on genuinely unrecoverable cases.

Sources

Drill Prompt Engineering & Structured Output

Practice only this domain’s questions, untimed, with instant explanations.