Domain 4 · Task 4.3
Structured Output with Tool Use
Enforce structured output using tool use and JSON schemas.
The most reliable way to get structured output from Claude is tool_use with a JSON schema: you define an extraction tool, the model fills its inputs, and you read the structured data straight from the tool_use response. This eliminates JSON syntax errors — no missing braces, no wrong types. It does not eliminate semantic errors (totals that don't add up, values in the wrong field, fabricated data). Getting this right is about two things: choosing the right tool_choice mode, and designing the schema so the model has a truthful way to represent missing or unclear data.
Key concept
Guaranteed structure → tool_use + schema. Unknown doc type with several schemas → tool_choice "any". Mandatory first extraction → force the named tool. Possibly-absent data → nullable/optional, never required.
What you need to know
tool_use + JSON schema eliminates syntax errors
tool_use with a JSON schema is the most reliable path to schema-compliant output because the model's response is validated against the schema's structure — guaranteeing syntactic validity (correct braces, types, no trailing commas). What it does not guarantee is semantic correctness: line items that don't sum to the stated total, a value placed in the wrong field, or a fabricated value all pass the schema. Those need validation and retry (Module 4.4). Read the extracted data from the tool_use block of the response, not from any accompanying text.
{
"name": "extract_metadata",
"description": "Extract structured metadata from a document.",
"input_schema": {
"type": "object",
"properties": {
"category": { "type": "string", "enum": ["bug", "feature", "docs", "unclear", "other"] },
"category_detail": { "type": ["string", "null"], "description": "Detail when category is 'other' or 'unclear'" },
"severity": { "type": "string", "enum": ["critical", "high", "medium", "low"] },
"ship_date": { "type": ["string", "null"], "description": "ISO 8601; null if absent in source" }
},
"required": ["category", "severity"]
}
}tool_choice: auto vs any vs forced
The tool_choice field controls whether and which tool the model must call:
"auto"— the model may call a tool or may return plain text. Use when a tool call is optional."any"— the model must call some tool, but chooses which. Use to guarantee structured output when you have several schemas and the document type is unknown — the model picks the matching extractor.- Forced —
{"type": "tool", "name": "extract_metadata"}makes the model call that specific tool. Use when a particular extraction must run first (a mandatory first pass).
The common exam move: "guarantee valid JSON" → tool_use + schema; "multiple possible schemas, unknown doc type" → "any"; "this extraction must run first" → force the named tool.
Schema design: nullable, optional, and enum + other
The schema is where you prevent fabrication:
- Required vs optional — mark a field
requiredonly if it is *always* present in the source. A required field pushes the model to fabricate when the data is missing. - Nullable (
"type": ["string", "null"]) — lets the model returnnullinstead of inventing an absent value. "unclear"enum value — better than forcing a confident-but-wrong category."other"+ a detail string — captures values outside your predefined set without data loss.
Include format-normalization rules in the prompt alongside the strict schema (dates → ISO 8601, currency → amount + code), because the schema enforces shape, not semantics.
Exam traps
| The trap | The reality |
|---|---|
tool_use + JSON schema guarantees the extracted data is correct. | It eliminates syntax errors only. Semantic errors (totals not summing, wrong field, fabrication) still pass the schema and need validation/retry. |
| To handle several possible document types, force one specific extraction tool. | When the doc type is unknown, use tool_choice "any" so the model picks the matching schema. Forcing a tool is for a mandatory first extraction. |
| Marking every field required makes extraction more complete. | Required fields push the model to fabricate when data is absent. Make possibly-absent fields nullable/optional instead. |
| A fixed enum with no escape hatch is the cleanest way to categorize. | Add "unclear" and "other" + a detail string so values outside the set are captured without data loss or a confident-but-wrong label. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
You're extracting structured data from medical notes into a schema with fields including diagnosis_codes (ICD-10), medications (with dosages), and follow_up_needed (boolean). Two persistent problems:
diagnosis_codessometimes contain free-text descriptions instead of codes ("hypertension" instead of "I10")follow_up_neededis set to true even when the note explicitly says "no follow-up needed"
Which combination best addresses both?
diagnosis_codes a regex-constrained string pattern; for follow_up_needed , replace the boolean with an enum ["yes", "no", "not_addressed"] and require a source_quote fieldCorrectWhy: Both problems are fundamentally schema/contract issues. ICD-10 codes have a strict format — a regex-constrained string pattern is exactly the contract the model should be handed. The boolean follow_up_needed is too permissive: forcing an enum that distinguishes explicit-no from not-addressed, plus requiring a source_quote , makes the model surface its evidence and reduces hallucinated booleans. Few-shots (A) help interpretation but don't constrain shape; regex post-processing for "no follow-up" misses paraphrases like "patient declined further visits." Temperature (C) doesn't define the right format. Two-step pipelines (D) add cost when the schema layer can solve it.
A junior engineer argues: "JSON mode and tool_use both produce structured output, so they're interchangeable — pick whichever feels easier."
Which statement is most accurate?
tool_use with an input_schema constrains the output to a specific schema; for reliable structured extraction, tool_use is preferredCorrecttool_use should be used in new codetool_use is exclusively for actions/side effects; using it for data extraction is a misuse — JSON mode is the correct choice for extractionWhy: JSON mode guarantees the output parses as JSON. tool_use with an input_schema constrains the output to a specific schema — field names, types, required fields, enums, regex patterns. For reliable structured extraction where downstream code depends on a precise shape, tool_use with a tightly-defined schema is the canonical pattern. JSON mode isn't deprecated (C). tool_use for extraction (D) is the recommended pattern — many Anthropic examples wrap extraction in a "structured_output" tool to enforce schema discipline.
Build exercise
Build a schema-enforced document extractor
~45 min- 1Define an
extract_metadatatool with a JSON schema: enum category (including "unclear" and "other"), a nullablecategory_detail, a severity enum, and a nullableship_date.Why:
tool_use+ schema is the most reliable path to syntactically valid output. - 2Run extraction on a document that is missing the ship date, and read the result from the
tool_useblock.You should see:
ship_datecomes back null rather than a fabricated date, because the field is nullable, not required. - 3Add a second extractor for a different document type and set
tool_choiceto "any" on a document whose type you don't specify.You should see: The model selects the matching extractor on its own.
- 4Switch to a forced
tool_choice({ type: "tool", name: "extract_metadata" }) and confirm that extraction always runs first.Why: Forcing a named tool is how you make a mandatory first-pass extraction deterministic.
- 5Feed a document with line items that don't sum to the stated total and confirm the schema still validates it.
You should see: The output is schema-valid despite the arithmetic error — proof that semantic checks belong in the next module.
Sources
Drill Prompt Engineering & Structured Output
Practice only this domain’s questions, untimed, with instant explanations.