CCAF logo

Domain 2 · Task 2.2

Structured Error Responses

Implement structured error responses for MCP tools.

When an MCP tool fails, *how* it fails determines whether the agent can do anything useful about it. MCP marks failures with the isError flag, but the flag alone is not enough. A structured, categorized error tells the agent whether to retry, fix its input, escalate, or explain a policy to the user. A uniform "Operation failed" string leaves the agent blind — it cannot tell a transient timeout from a permanent permission denial, so it either wastes retries or gives up when it shouldn't.

Key concept

Structured, categorized errors enable intelligent recovery; generic errors destroy it. Return an errorCategory, an isRetryable boolean, a human-readable message, and — for multi-agent flows — the attempted query and any partial results.

What you need to know

The isError flag and the four categories

MCP tools signal failure by returning isError: true in the result (distinct from a JSON-RPC protocol error). What makes the error *actionable* is the metadata alongside it. Classify every failure into one of four categories, each with a different recovery path:

CategoryRetryable?Agent action
Transient (timeout, 503)YesRetry with backoff locally (1–2x), then propagate
Validation (bad input)No (fix input)Modify the request and retry
Business (policy/threshold)NoExplain to the user; offer an alternative
Permission (access denied)NoEscalate

The retryable vs non-retryable distinction is the point: structured metadata stops the agent from wasting retries on a validation or permission error that will never succeed on repeat.

Shape of a structured error

Return, at minimum: errorCategory (transient / validation / permission / business), an isRetryable boolean, and a human-readable message. For business-rule violations, include retriable: false plus a customer-friendly explanation so the agent can relay *why* something is disallowed rather than just "failed". Contrast the two shapes below — the structured one gives the agent something to reason about; the generic one gives it nothing.

json
// Structured (good) — the agent can decide to retry, re-query, or escalate:
{
  "isError": true,
  "content": {
    "errorCategory": "transient",
    "isRetryable": true,
    "message": "Orders API timed out.",
    "attempted_query": "order_id=12345",
    "partial_results": null
  }
}

// Generic (anti-pattern) — nothing to reason about:
{ "isError": true, "content": "Operation failed" }

Local recovery, then propagate with context

In multi-agent workflows, handle what you can locally and propagate only what you can't. For transient failures, retry inside the subagent 1–2 times with backoff before giving up — do not run infinite retries. If it still fails, propagate to the coordinator with the failure type, the attempted query, any partial results, and suggested alternatives — never a bare status. This lets the coordinator continue the overall workflow with partial coverage instead of terminating it because one subagent failed.

python
def search_with_recovery(query):
    for attempt in range(2):  # local recovery: 1-2 tries, NOT infinite
        result = search(query)
        if not result.get("isError"):
            return result
        if not result["content"].get("isRetryable"):
            break  # validation/permission/business won't fix on retry
        backoff(attempt)
    # propagate structured context to the coordinator, not a bare status
    return {
        "isError": True,
        "content": {
            "errorCategory": result["content"]["errorCategory"],
            "isRetryable": False,
            "message": "Search subagent could not complete.",
            "attempted_query": query,
            "partial_results": result["content"].get("partial_results"),
            "alternatives": ["retry with a narrower date range", "search cached index"],
        },
    }

Access failure is not an empty result

A dangerous silent bug is treating an access failure as a valid empty result. "The search returned nothing" and "the search never ran because it was denied/timed out" look identical if you only check for an empty list — but they demand opposite responses. A valid empty result is a successful query with no matches (report it and move on); an access failure needs a retry/escalation decision. Always distinguish them: mark failures with isError and category, and never let a suppressed error masquerade as a successful empty query.

Exam traps

The trapThe reality
When a tool fails, returning { "isError": true, "content": "Operation failed" } is a reasonable, safe default.A uniform "Operation failed" prevents appropriate recovery — the agent can't tell transient from permanent. Return category + retryable + message + attempted query.
If a search subagent times out, returning empty results marked as success keeps the workflow moving.That silently suppresses a failure and is indistinguishable from a valid empty result. Distinguish access failure from a genuine no-match, and surface the failure.
One subagent failing means the whole multi-agent workflow should terminate.Continue with partial results and annotate the coverage gap. Propagate structured context to the coordinator so it can proceed, not kill the run.
A transient timeout should be retried until it succeeds.Infinite retries are an anti-pattern. Do local recovery 1–2 times with backoff, then propagate a structured error upward.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your MCP tool for booking calendar events currently returns errors as plain strings: "Failed", "Could not book", "Conflict detected." In production, when these errors occur, Claude variously retries indefinitely, gives up immediately, or asks the user nonsensical clarifying questions. You want to improve error handling without rewriting the agent's prompt.

Most effective change?

AReturn errors as structured objects with errorCode (enum), retryable (boolean), userMessage (string), developerMessage (string)Correct
BAlways retry failed tool calls with exponential backoff at the MCP server layer before returning to the agent
CReturn errors in a consistent format like "ERROR_TYPE: human-readable description" so Claude can text-parse them reliably
DHave the MCP tool return a 200 OK with status: "error" in the body so Claude doesn't treat them as exceptions

Why: Structured error responses are the standard MCP pattern for this exact problem. errorCode lets the model recognize the category; retryable: false deterministically tells it not to retry; userMessage gives it ready language to surface; developerMessage aids debugging. Option B (server-side retry only) hides retry logic from the agent and doesn't help on permanent errors. Option C asks the model to text-parse error strings — exactly the unreliable behavior you're trying to fix. Option D misuses HTTP status semantics and breaks tooling that expects errors to be visible to the agent.

Your structured-extraction pipeline uses an MCP tool that calls a downstream knowledge base. When the knowledge base is temporarily unavailable, the tool returns an error. When the queried record genuinely doesn't exist in the knowledge base, the tool currently also returns an error with the same "not found" message. Your agent retries identically in both cases, wasting calls when records genuinely don't exist.

What's the correct fix?

AReturn the knowledge base error with isError: true for service failures, and return an empty result set marked successful for genuine not-found cases.
BDistinguish the two cases explicitly: return the service failure with isError: true and isRetryable: true ; return the not-found case as a valid empty result (no isError ) so the agent treats it as a successful query with zero matches.Correct
CWrap both cases in isError: true but include a structured errorCategory field distinguishing "service_unavailable" from "not_found" so the agent can decide whether to retry.
DImplement local retry logic inside the MCP tool for service failures, only surfacing errors to the agent after all retries are exhausted.

Why: Task Statement 2.2 makes this distinction explicit: distinguishing access failures from valid empty results in error reporting. Service failures = isError: true with isRetryable: true ; genuine not-found = a valid empty result (no isError ), representing a successful query with zero matches. Option A is on the right track but mixes the concept — isError: true for service failures is correct, but a "successful empty result" doesn't need an isError field at all. Option C keeps both under isError which conflates two semantically different cases. Option D removes the agent's visibility into transient issues entirely.

Build exercise

Add structured errors to a flaky MCP tool

~45 min
  1. 1
    Take a tool that currently returns { isError: true, content: "Operation failed" } and replace the content with errorCategory, isRetryable, and a human-readable message.

    Why: The category + retryable pair is what lets the agent choose a recovery path instead of guessing.

  2. 2
    Simulate all four categories (timeout → transient, bad input → validation, policy block → business, 403 → permission) and confirm each returns the right isRetryable value.

    You should see: Only transient is retryable; validation/business/permission are false with distinct recovery actions.

  3. 3
    For a business-rule violation, add retriable: false plus a customer-friendly explanation of why the action is disallowed.

    Why: The agent should be able to explain the policy and offer an alternative, not just say it failed.

  4. 4
    Wrap the tool in a subagent that retries transient errors 1–2 times with backoff, then propagates a structured error (attempted query + partial results + alternatives) to the coordinator.

    Why: Local recovery contains transient blips; propagation with context keeps the coordinator able to continue.

    You should see: The coordinator receives enough context to continue with partial coverage rather than aborting.

  5. 5
    Add a test that a denied query and a genuinely empty result produce different responses (one flagged isError, one a successful no-match).

    Why: Proves you are distinguishing access failure from a valid empty result — the silent-suppression bug.

Sources

Drill Tool Design & MCP Integration

Practice only this domain’s questions, untimed, with instant explanations.