Domain 4 · Task 4.4
Validation, Retry & Feedback Loops
Implement validation, retry, and feedback loops for extraction quality.
tool_use gives you syntactically valid output; validation and retry give you semantically correct output. When an extraction fails a check — line items don't sum, a value is in the wrong field — the fix is retry with error feedback: send the model the original document, its failed extraction, and the specific error, and ask it to self-correct. The critical judgment is knowing the limit of retry: it fixes format, structural, and arithmetic errors, but it cannot recover information that is simply not present in the source.
Key concept
Format / structure / arithmetic error → retry with the specific error and the original document. Information absent from the source → retry won't help; get the missing source or route to a human.
What you need to know
Retry with error feedback
The retry loop re-prompts with three things: the original input + the failed output + the specific validation error. Naming the exact error is what makes self-correction work — "the line items sum to $145.00 but stated_total is $150.00" guides the correction far better than "try again." This is effective for semantic errors that survive the schema:
- Values that don't add up (arithmetic).
- A value placed in the wrong field (structural).
- Format mismatches (a date that isn't ISO 8601).
Syntax errors don't need this loop — tool_use already eliminates them.
def extract_with_retry(doc, max_retries=2):
result = extract(doc)
for _ in range(max_retries):
errors = validate(result) # semantic checks: sums, fields, formats
if not errors:
return result
if is_absent_from_source(errors): # retry cannot conjure missing data
return route_to_human(doc, errors)
result = extract(
doc,
prior_output=result,
feedback=f"Validation failed: {errors}. Re-extract, correcting these.",
)
return route_to_human(doc, errors)The limit of retry: absent information
Retry is powerful but not omnipotent. If a required value is absent from the source — for instance the ship date lives in a separate document you never supplied — no amount of re-prompting will produce it, because the model has nothing to read. Distinguish the two cases:
- Retry will help: the data is in the document but was mis-parsed, mis-placed, or mis-formatted.
- Retry won't help: the data is genuinely not in the provided source. The fix is to supply the missing source or route to a human — not to loop.
A loop that keeps retrying on absent data either spins forever or, worse, pressures the model into fabricating the value.
Self-correction and feedback-loop design
You can make discrepancies surface deterministically by extracting a stated value alongside a computed one and a conflict flag, so validation is a simple boolean check instead of re-derivation. Add conflict_detected booleans for inconsistent sources. For feedback-loop design over time, add a detected_pattern field that records which construct triggered each finding — then, when developers dismiss findings, you can analyze the dismissal patterns and tighten the criteria for the categories that produce false positives.
{
"stated_total": "$150.00",
"calculated_total": "$145.00",
"conflict_detected": true,
"detected_pattern": "line_items_sum_mismatch"
}Exam traps
| The trap | The reality |
|---|---|
| Extraction failed validation because a field is missing — retry with feedback will recover it. | If the field is absent from the provided source (e.g. it lives in an unsupplied document), retry can't conjure it. Get the source or route to a human. |
| Retrying with a generic "that was wrong, try again" is enough. | Effective retry sends the original document, the failed output, and the specific error. Naming the exact error is what drives correction. |
tool_use already validates the data, so a separate validation step is redundant. | tool_use eliminates syntax errors only. Semantic checks (sums, field placement, formats) are a separate, necessary step. |
| To catch a total that doesn't match, re-derive it in downstream code. | Have the model emit stated_total and calculated_total with a conflict_detected flag so the discrepancy surfaces deterministically at extraction time. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
Your structured-extraction pipeline uses tool_use with a strict JSON schema. After deployment, JSON syntax errors dropped to 0% — the schema enforces structure. But you see a new class of error: values that match the schema's type but are semantically wrong. Examples: a date field receives "2024-13-45" (string matches pattern but isn't a real date); a tax_rate field receives 0.85 (correct type, but 85% — implausibly high for any tax jurisdiction in your dataset).
What addresses this?
tool_use to JSON mode with a more permissive schemaWhy: tool_use with a JSON schema guarantees syntactic compliance (output will parse and field types match) but does not guarantee semantic correctness. "2024-13-45" matches a string field's type; 0.85 matches a number's type. Programmatic validation on top — real date parsing, plausibility checks on numeric ranges, business-rule validation — is the layer that catches semantic errors. Few-shots (D) help at the margin but aren't a guarantee. Temperature (B) doesn't change what the model considers semantically valid. Switching to JSON mode (A) is strictly worse than tool_use for reliability.
Your data extraction pipeline processes restaurant menus. The schema includes prices: number[] . After deployment, ~3% of extractions fail validation: prices come back as ["$12.99", "$15.50"] (strings with dollar signs) instead of [12.99, 15.50] (numbers). You implement a retry-with-error-feedback approach where on validation failure, you re-prompt the model with the source document, the failed extraction, and the specific error. Most failures resolve within 2-3 attempts. For which failure pattern would additional retries be LEAST likely to resolve the issue?
Why: Retries can't add information that isn't in the source. This is the documented Task 4.4 limit. A, B, C are all structural/format mismatches where the model has the right value but wrong shape — retry-with-feedback handles those.
Build exercise
Build a validation-and-retry extraction loop
~50 min- 1Extract an invoice with
tool_use, then write a validator that checks line items sum to the stated total and dates are ISO 8601.Why: Semantic validation catches the errors that
tool_uselets through. - 2On failure, re-prompt with the original document, the failed extraction, and the exact error string.
You should see: The model self-corrects an arithmetic or format mistake on the next pass.
- 3Add a
stated_total/calculated_total/conflict_detectedpattern so mismatches surface as a boolean.Why: Deterministic conflict flags turn validation into a simple check instead of re-derivation.
- 4Feed a document whose required
ship_dategenuinely lives in a separate file you don't provide, and detect that retry doesn't resolve it.You should see: The value stays null across retries — the signal to stop looping and route elsewhere.
- 5Add a branch that routes the absent-data case to a human (or requests the missing source) instead of retrying.
Why: Retry can't conjure missing information; the correct escalation is to get the source or a human.
- 6Cap retries (e.g. 2) so a persistently-failing extraction escalates rather than looping forever.
Why: A cap prevents runaway loops and forces escalation on genuinely unrecoverable cases.
Sources
Drill Prompt Engineering & Structured Output
Practice only this domain’s questions, untimed, with instant explanations.