CCAF logo

Domain 5 · Task 5.3

Error Propagation in Multi-Agent Systems

Implement error-propagation strategies across multi-agent systems.

In a coordinator–subagent system, a subagent failure is not the end of the workflow — it is information the coordinator needs to route around. The difference between a resilient system and a brittle one is what a failing subagent reports. This lesson covers structured error context, the crucial distinction between an access failure and a valid empty result, and the two anti-patterns — silent suppression and whole-workflow termination — that the exam flags on sight.

Key concept

Propagate structured context plus partial results; never suppress silently, and never kill the whole workflow on one failure.

What you need to know

Structured error context

A generic status ("Operation failed", "search unavailable") hides everything the coordinator needs to make a recovery decision. Instead, a failing subagent should return structured error context: the failure type, the attempted query, any partial results it did gather, and alternative approaches to try. With that, the coordinator can retry a timeout, narrow a query, or continue with what's covered — an intelligent decision the generic status makes impossible.

json
{
  "status": "partial_failure",
  "failure_type": "timeout",
  "attempted_query": "AI impact on music 2024",
  "partial_results": [
    { "title": "...", "url": "...", "relevance": 0.8 }
  ],
  "alternative_approaches": ["Narrower query: 'AI music composition tools'"],
  "coverage_impact": "Music production not covered"
}

Access failure vs valid empty result

The single most important distinction is between an access failure and a valid empty result. A timeout, a 500, or a dropped connection is an access failure — the query *might* have matched, so the coordinator should weigh a retry. Zero rows from a search that ran correctly is a valid empty result — a genuine success meaning "no matches", which must not be retried and must not be reported as a failure. Collapsing the two (empty = success, or timeout = empty) is the root of silent-suppression bugs.

Local recovery, then propagate

Subagents should attempt local recovery for transient failures — a bounded 1–2 retries — and propagate only what stays unresolved, carrying what was attempted and any partial results. Two extremes are wrong: infinite retries inside a subagent (it spins forever, never surfacing the problem) and immediate propagation of every blip (it floods the coordinator). At synthesis time, structure the output with coverage annotations that mark which findings are well-supported versus which have gaps because a source was unavailable.

The two anti-patterns

Two failure modes are near-automatic wrong answers:

  • Silent suppression — treating an empty or errored result as success so the failure never surfaces. The workflow completes "successfully" on incomplete data, which is worse than a visible error.
  • Terminating the whole workflow on one failure — one subagent's timeout should not abort a research task with three other healthy branches. Continue with partial results and annotate the coverage gap; let the coordinator decide whether the remaining coverage is sufficient.

Exam traps

The trapThe reality
A failing subagent should return a clean generic status like "search unavailable" so the coordinator stays simple.Generic statuses strip the context needed to recover. Return failure type, attempted query, partial results, and alternatives.
Zero results and a timeout can both be treated as "no data" — either way there's nothing to work with.A valid empty result is a success (no matches); an access failure might have matched and may warrant a retry. Conflating them causes silent data loss.
If one subagent fails, terminate the workflow so you don't report partial or misleading results.Continue with partial results and annotate the coverage gap. One branch's failure shouldn't discard the work of the others.
Retrying inside the subagent until it succeeds is the most robust recovery.Bounded local recovery (1–2 retries), then propagate. Unbounded retries spin forever and hide the failure from the coordinator.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

A research subagent in your pipeline times out after 90 seconds while researching "quantum computing investments." The coordinator was expecting the response within 60 seconds and has begun rerunning the same query, then again, then again. After 4 retries the coordinator gives up and returns "research failed" to the user, even though the subagent had partial results before each timeout. Your team needs to redesign this error propagation.

AIncrease the coordinator's timeout to 5 minutes so the subagent has time to complete.
BImplement exponential backoff with a max of 3 retries.
CHave the subagent return structured error context to the coordinator including the failure type ("timeout"), the attempted query, any partial results gathered before timeout, and potential alternative approaches ("try narrower query about quantum computing patents instead").Correct
DHave the subagent automatically retry internally with a narrower query if the broad query times out.

Why: Structured error context including partial results and alternatives (Task 5.3) enables intelligent coordinator recovery. Timeouts (A) shouldn't be hidden by simple extension. Backoff retries (B) don't help if the query itself is too broad. Internal retries with narrower queries (D) take the coordinator out of the recovery decision.

The document analysis subagent encounters a corrupted PDF file it cannot parse. When designing the system's error handling, what is the most effective way to handle this failure?

AReturn the error with context to the coordinator agent, letting it decide how to proceed.Correct
BSilently skip the corrupted document and continue processing other files to avoid interrupting the workflow.
CThrow an exception that terminates the entire research workflow.
DAutomatically retry parsing the document three times with exponential backoff before reporting failure.

Why: Returning errors with context to the coordinator is the correct approach because it allows the coordinator agent to make informed decisions about how to proceed—whether to skip the document, try an alternative, or inform the user. Silent failures (B) hide critical information and undermine report completeness. Terminating the entire workflow (C) is disproportionate for a single file issue. Retrying (D) is pointless for corrupted files since corruption is a permanent failure, not a transient one that might resolve.

Build exercise

Add structured error propagation to a research coordinator

~45 min
  1. 1
    Build a coordinator that fans out to three search subagents, one of which you can force to time out and one whose query legitimately returns zero results.

    Why: You need both a real access failure and a valid empty result to test that they're handled differently.

  2. 2
    Start with a baseline where a failing subagent returns just "failed" and the coordinator aborts the whole run.

    You should see: One timeout discards the two healthy results — the whole-workflow-termination anti-pattern.

  3. 3
    Change the subagent to return a structured error object with failure_type, attempted_query, partial_results, and alternative_approaches.

    Why: Structured context is what lets the coordinator recover instead of aborting.

  4. 4
    Have the subagent attempt 1–2 local retries on a timeout before propagating, and confirm it stops retrying after the cap.

    You should see: Transient timeouts self-heal; persistent ones propagate with their partial results rather than spinning.

  5. 5
    Make the coordinator distinguish the empty-result subagent (success, no matches) from the timed-out one (retry or route around).

    You should see: The empty result is recorded as coverage, not an error, and is never retried.

  6. 6
    Have the coordinator synthesize a final report that continues with partial results and annotates the coverage gap from the unavailable source.

    Why: Coverage annotations make the gap visible instead of silently suppressed.

Sources

Drill Context Management & Reliability

Practice only this domain’s questions, untimed, with instant explanations.