Domain 2 · Task 2.1
Tool Interface Design
Design effective tool interfaces with clear descriptions and boundaries.
A tool's description is the main thing the model reads to decide whether to call it. When two tools have thin or overlapping descriptions, Claude cannot tell them apart and calls the wrong one. The lowest-effort, highest-leverage fix for almost every tool-selection bug is therefore to rewrite the descriptions (and check the system prompt) — not to bolt on a routing layer or pile on examples.
Key concept
Selection problem → fix descriptions (and check the system prompt) first. It is the lowest-effort, highest-leverage, root-cause fix; renaming and splitting tools are the follow-ups when descriptions alone can't disambiguate.
What you need to know
Descriptions drive selection
Tool descriptions are the primary mechanism the model uses for tool selection. A minimal description like "Analyzes content" gives Claude nothing to distinguish it from a sibling tool, so selection becomes unreliable. A good description spells out:
- What the tool does and, crucially, when to use it vs a similar alternative.
- Input formats — the exact shape and type of each argument.
- Example queries the tool is meant to handle.
- Edge cases and boundary explanations — what it does *not* cover.
Ambiguous or overlapping descriptions (analyze_content vs analyze_document) cause misrouting: the model picks whichever name it pattern-matched first. Rich, differentiated descriptions are what make selection reliable.
{
"name": "extract_web_results",
"description": "Extract structured results from a WEB SEARCH response (a list of URLs + snippets). Use this ONLY for live web content. Do NOT use for uploaded files or database rows — use extract_data_points for those. Input: the raw search-results JSON. Returns: title, url, and a one-line summary per result.",
"input_schema": {
"type": "object",
"properties": {
"search_results": { "type": "string", "description": "Raw JSON returned by the web_search tool." }
},
"required": ["search_results"]
}
}Rename to remove overlap
When two names invite confusion, rename so each tool's purpose is unmistakable. A generic analyze_content that only ever handles web output should become extract_web_results with a web-specific description. The rename does double duty: it eliminates the name collision *and* forces you to write a purpose-specific description. The goal is that each tool's name and description together answer "when would I pick this over the others?" without ambiguity.
Split generic tools into purpose-specific ones
A single overloaded tool that "does everything to a document" is hard for the model to aim. Split it into narrow, single-purpose tools. For example, one analyze_document becomes:
extract_data_points— pull specific fields/figures.summarize_content— produce a summary.verify_claim_against_source— check a claim against the source text.
Each resulting tool has one job, one clear description, and one obvious trigger — so the model routes correctly by construction. This also composes well with least-privilege tool distribution (Task 2.3): narrow tools are easier to scope per agent.
System-prompt wording can override descriptions
Tool selection is not decided by descriptions alone. System-prompt wording can create unintended tool associations that override what a description says. If the system prompt repeatedly ties a keyword (say, "customer") to a concept, the model may reach for get_customer even on an order query, regardless of that tool's description. So when you debug a selection problem, review the system prompt for keyword-sensitive instructions as part of the same fix — don't assume the description is the only input.
Exam traps
| The trap | The reality |
|---|---|
| The agent keeps calling the wrong tool, so add few-shot examples of correct tool calls to the prompt. | Few-shot examples add tokens without fixing the root cause. The cause is thin/overlapping descriptions — expand and differentiate the descriptions first. |
| Misrouting between similar tools means you need a routing/classifier layer in front of them. | A routing layer is over-engineered for a selection bug. Descriptions are the model's routing mechanism — fix them (and the system prompt) before adding infrastructure. |
| When two tools overlap, consolidating them into one tool is the cleanest first step. | Consolidation is more effort than a "first step" warrants and can create an overloaded tool. Prefer rewriting/renaming; split (not merge) when a tool is too generic. |
| The tool description is right, so the system prompt can't be the reason the model picks the wrong tool. | System-prompt wording can create keyword associations that override descriptions. Reviewing the system prompt is part of the selection fix. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
You have an MCP server with a single execute_database_query tool that takes an arbitrary SQL string. In production, Claude sometimes runs destructive queries (DELETE, DROP) when the user only asked to investigate data. The tool description is detailed and warns against destructive operations. Users do legitimately need destructive operations for other tasks.
The most effective fix is:
read_query (SELECT only) and write_query (INSERT/UPDATE/DELETE/DDL), with write_query requiring an explicit confirmation step in its contractCorrectWhy: This is a tool-design problem, not a description problem. A single tool that can do anything from "show me data" to "drop the database" forces the model to make a high-stakes safety judgment on every call. Splitting into read_query and write_query gives an explicit, well- typed safety boundary — Claude picks the read tool for investigation, the write tool for mutations. A hook (C) would block destructive queries entirely, but the question states users legitimately need them. Description improvements (A) and prompt warnings (D) leave compliance to model discretion. Tool granularity matched to user intent is the root-cause fix.
You've added an MCP server with a tool called extract_pdf_text . The agent has a built-in Read tool that also handles PDFs. In testing, the agent uses Read for PDFs (which works but loses table structure) instead of extract_pdf_text (which preserves tables). Both tools are listed and the MCP server is healthy.
Most effective fix?
extract_pdf_text 's description to explicitly mention it preserves table structure (which Read does not) and clarify when to prefer it over ReadCorrectextract_pdf_text over Read."extract_pdf_text first; if it fails, fall back to ReadWhy: Tool selection is driven by the descriptions Claude sees. When extract_pdf_text has a generic description and Read has rich, well-documented behavior, the agent picks Read. Beefing up the MCP tool's description — explicitly mentioning table preservation as an advantage, clarifying when to prefer it — is the root-cause fix. Removing Read (A) breaks legitimate uses elsewhere. Prompt-level routing (C) is weaker than tool descriptions for selection. Try-then-fallback (D) is a fragile pattern that wastes calls. This is one of the most under-appreciated MCP lessons: descriptions are the selection contract.
Build exercise
Diagnose and fix a tool-selection bug
~40 min- 1Define two tools with deliberately minimal descriptions — e.g.
get_customer("Gets customer info") andget_order("Gets order info") — and ask the agent an order-status question.Why: You need a reproducible misrouting bug before you can prove which fix actually removes it.
You should see: The model sometimes calls
get_customerfor an order query because neither description tells it where the boundary is. - 2Rewrite each description to state purpose, inputs, outputs, and an explicit "use this instead of the other when…" boundary. Re-run the same query.
Why: Descriptions are the primary selection signal; differentiating them is the root-cause fix.
You should see: The model now routes order queries to
get_orderreliably. - 3Add a generic
analyze_documenttool and confirm the model struggles to aim it, then split it intoextract_data_points,summarize_content, andverify_claim_against_source.Why: Single-purpose tools route correctly by construction; a generic one forces the model to guess intent.
- 4Rename an overlapping tool (e.g.
analyze_content→extract_web_results) and give it a web-specific description.Why: Renaming removes the name collision and forces a purpose-specific description in one move.
- 5Introduce a system-prompt line that over-associates a keyword ("always start with the customer record") and observe it override a correct description; then remove/soften it.
Why: Proves that system-prompt wording is part of the selection surface, not just the descriptions.
You should see: Selection flips back to correct once the keyword-sensitive instruction is removed.
Sources
Drill Tool Design & MCP Integration
Practice only this domain’s questions, untimed, with instant explanations.