Domain 2 · Task 2.2
Structured Error Responses
Implement structured error responses for MCP tools.
When an MCP tool fails, *how* it fails determines whether the agent can do anything useful about it. MCP marks failures with the isError flag, but the flag alone is not enough. A structured, categorized error tells the agent whether to retry, fix its input, escalate, or explain a policy to the user. A uniform "Operation failed" string leaves the agent blind — it cannot tell a transient timeout from a permanent permission denial, so it either wastes retries or gives up when it shouldn't.
Key concept
Structured, categorized errors enable intelligent recovery; generic errors destroy it. Return an errorCategory, an isRetryable boolean, a human-readable message, and — for multi-agent flows — the attempted query and any partial results.
What you need to know
The isError flag and the four categories
MCP tools signal failure by returning isError: true in the result (distinct from a JSON-RPC protocol error). What makes the error *actionable* is the metadata alongside it. Classify every failure into one of four categories, each with a different recovery path:
| Category | Retryable? | Agent action |
|---|---|---|
| Transient (timeout, 503) | Yes | Retry with backoff locally (1–2x), then propagate |
| Validation (bad input) | No (fix input) | Modify the request and retry |
| Business (policy/threshold) | No | Explain to the user; offer an alternative |
| Permission (access denied) | No | Escalate |
The retryable vs non-retryable distinction is the point: structured metadata stops the agent from wasting retries on a validation or permission error that will never succeed on repeat.
Shape of a structured error
Return, at minimum: errorCategory (transient / validation / permission / business), an isRetryable boolean, and a human-readable message. For business-rule violations, include retriable: false plus a customer-friendly explanation so the agent can relay *why* something is disallowed rather than just "failed". Contrast the two shapes below — the structured one gives the agent something to reason about; the generic one gives it nothing.
// Structured (good) — the agent can decide to retry, re-query, or escalate:
{
"isError": true,
"content": {
"errorCategory": "transient",
"isRetryable": true,
"message": "Orders API timed out.",
"attempted_query": "order_id=12345",
"partial_results": null
}
}
// Generic (anti-pattern) — nothing to reason about:
{ "isError": true, "content": "Operation failed" }Local recovery, then propagate with context
In multi-agent workflows, handle what you can locally and propagate only what you can't. For transient failures, retry inside the subagent 1–2 times with backoff before giving up — do not run infinite retries. If it still fails, propagate to the coordinator with the failure type, the attempted query, any partial results, and suggested alternatives — never a bare status. This lets the coordinator continue the overall workflow with partial coverage instead of terminating it because one subagent failed.
def search_with_recovery(query):
for attempt in range(2): # local recovery: 1-2 tries, NOT infinite
result = search(query)
if not result.get("isError"):
return result
if not result["content"].get("isRetryable"):
break # validation/permission/business won't fix on retry
backoff(attempt)
# propagate structured context to the coordinator, not a bare status
return {
"isError": True,
"content": {
"errorCategory": result["content"]["errorCategory"],
"isRetryable": False,
"message": "Search subagent could not complete.",
"attempted_query": query,
"partial_results": result["content"].get("partial_results"),
"alternatives": ["retry with a narrower date range", "search cached index"],
},
}Access failure is not an empty result
A dangerous silent bug is treating an access failure as a valid empty result. "The search returned nothing" and "the search never ran because it was denied/timed out" look identical if you only check for an empty list — but they demand opposite responses. A valid empty result is a successful query with no matches (report it and move on); an access failure needs a retry/escalation decision. Always distinguish them: mark failures with isError and category, and never let a suppressed error masquerade as a successful empty query.
Exam traps
| The trap | The reality |
|---|---|
When a tool fails, returning { "isError": true, "content": "Operation failed" } is a reasonable, safe default. | A uniform "Operation failed" prevents appropriate recovery — the agent can't tell transient from permanent. Return category + retryable + message + attempted query. |
| If a search subagent times out, returning empty results marked as success keeps the workflow moving. | That silently suppresses a failure and is indistinguishable from a valid empty result. Distinguish access failure from a genuine no-match, and surface the failure. |
| One subagent failing means the whole multi-agent workflow should terminate. | Continue with partial results and annotate the coverage gap. Propagate structured context to the coordinator so it can proceed, not kill the run. |
| A transient timeout should be retried until it succeeds. | Infinite retries are an anti-pattern. Do local recovery 1–2 times with backoff, then propagate a structured error upward. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
Your MCP tool for booking calendar events currently returns errors as plain strings: "Failed", "Could not book", "Conflict detected." In production, when these errors occur, Claude variously retries indefinitely, gives up immediately, or asks the user nonsensical clarifying questions. You want to improve error handling without rewriting the agent's prompt.
Most effective change?
Why: Structured error responses are the standard MCP pattern for this exact problem. errorCode lets the model recognize the category; retryable: false deterministically tells it not to retry; userMessage gives it ready language to surface; developerMessage aids debugging. Option B (server-side retry only) hides retry logic from the agent and doesn't help on permanent errors. Option C asks the model to text-parse error strings — exactly the unreliable behavior you're trying to fix. Option D misuses HTTP status semantics and breaks tooling that expects errors to be visible to the agent.
Your structured-extraction pipeline uses an MCP tool that calls a downstream knowledge base. When the knowledge base is temporarily unavailable, the tool returns an error. When the queried record genuinely doesn't exist in the knowledge base, the tool currently also returns an error with the same "not found" message. Your agent retries identically in both cases, wasting calls when records genuinely don't exist.
What's the correct fix?
service_unavailable" from "not_found" so the agent can decide whether to retry.Why: Task Statement 2.2 makes this distinction explicit: distinguishing access failures from valid empty results in error reporting. Service failures = isError: true with isRetryable: true ; genuine not-found = a valid empty result (no isError ), representing a successful query with zero matches. Option A is on the right track but mixes the concept — isError: true for service failures is correct, but a "successful empty result" doesn't need an isError field at all. Option C keeps both under isError which conflates two semantically different cases. Option D removes the agent's visibility into transient issues entirely.
Build exercise
Add structured errors to a flaky MCP tool
~45 min- 1Take a tool that currently returns
{ isError: true, content: "Operation failed" }and replace the content witherrorCategory,isRetryable, and a human-readablemessage.Why: The category + retryable pair is what lets the agent choose a recovery path instead of guessing.
- 2Simulate all four categories (timeout → transient, bad input → validation, policy block → business, 403 → permission) and confirm each returns the right
isRetryablevalue.You should see: Only
transientis retryable; validation/business/permission arefalsewith distinct recovery actions. - 3For a business-rule violation, add
retriable: falseplus a customer-friendly explanation of why the action is disallowed.Why: The agent should be able to explain the policy and offer an alternative, not just say it failed.
- 4Wrap the tool in a subagent that retries transient errors 1–2 times with backoff, then propagates a structured error (attempted query + partial results + alternatives) to the coordinator.
Why: Local recovery contains transient blips; propagation with context keeps the coordinator able to continue.
You should see: The coordinator receives enough context to continue with partial coverage rather than aborting.
- 5Add a test that a denied query and a genuinely empty result produce different responses (one flagged
isError, one a successful no-match).Why: Proves you are distinguishing access failure from a valid empty result — the silent-suppression bug.
Sources
Drill Tool Design & MCP Integration
Practice only this domain’s questions, untimed, with instant explanations.