CCAF logo

Domain 1 · Task 1.2

Multi-Agent Orchestration

Orchestrate multi-agent systems (coordinator–subagent).

A multi-agent system uses a coordinator (the hub) that delegates work to subagents (the spokes). The coordinator manages *all* inter-subagent communication, error handling, and routing; subagents run with isolated context and never automatically inherit the coordinator's history. Most exam questions here test one discrimination: when the output is incomplete, is the fault in a subagent or in how the coordinator split up the work?

Key concept

The coordinator owns scope and routing. If coverage is incomplete and the logs show narrow subtask assignment, the fix is at the coordinator's decomposition — not at the subagents, which executed their assigned scope correctly.

What you need to know

Hub-and-spoke: the coordinator owns everything

In a hub-and-spoke topology the coordinator is the single point that manages inter-subagent communication, error handling, and routing. It owns four responsibilities:

  • Decomposition — breaking the query into subtasks that together cover the whole scope.
  • Delegation — dispatching each subtask to a subagent with the context it needs.
  • Aggregation — synthesizing subagent outputs into one answer.
  • Dynamic selection — choosing *which* subagents to run based on query complexity, rather than always firing the full pipeline.

Routing all subagent communication through the coordinator is what gives you observability and consistent error handling. Subagents talking directly to each other is an anti-pattern: it hides what happened and scatters error handling.

Dynamic selection and scope partitioning

A good coordinator analyzes the query and dynamically selects subagents instead of always running the entire pipeline. It also partitions scope across subagents — distinct subtopics or source types — to minimize duplication of work. When two subagents research overlapping ground, you pay twice and still risk gaps; when the partition is too narrow, you get incomplete coverage of a broad topic. Scope design is the coordinator's most consequential decision.

Iterative refinement loops

Coverage rarely comes right on the first pass. The coordinator should run an iterative refinement loop: synthesize the current results, detect coverage gaps, re-delegate targeted searches or analyses for just those gaps, then re-synthesize — repeating until coverage is adequate. This keeps the coordinator, not the subagents, in charge of deciding when the answer is complete.

text
synthesize -> detect coverage gaps -> coordinator re-delegates targeted
search/analysis for the gaps -> re-synthesize -> repeat until coverage adequate

Diagnosing incomplete coverage

The canonical failure: a query about *creative industries* comes back covering only *visual arts*. It is tempting to blame the search, analysis, or synthesis subagents — but if the logs show each subagent executed its assigned scope correctly, the root cause is that the coordinator's decomposition was too narrow. Over-provisioning or re-running a downstream subagent does not fix a scope problem created upstream. Widen the coordinator's decomposition instead.

Exam traps

The trapThe reality
Subagents can message each other directly to hand off intermediate results.Direct subagent-to-subagent communication destroys observability and scatters error handling. Route all communication through the coordinator.
The coordinator should always run the full subagent pipeline so nothing is missed.Always running everything wastes work and duplicates coverage. Dynamically select subagents by query complexity.
Incomplete topic coverage means a subagent underperformed, so provision it more resources.If each subagent covered its assigned scope, the fault is the coordinator's narrow decomposition. Fix the coordinator's scope, not the subagent.
Once subagents return, the coordinator just concatenates their outputs and is done.The coordinator should evaluate the synthesis for gaps and re-delegate targeted queries until coverage is adequate — an iterative refinement loop.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

You're scaling a multi-agent legal research system. The current architecture: coordinator → research subagent → case-law subagent → synthesis subagent, where each subagent's output is forwarded by the coordinator to the next. A senior engineer proposes letting subagents call each other directly to reduce coordinator overhead — for example, the research subagent could call the case-law subagent when it encounters citations, without round-tripping through the coordinator.

What's the most accurate assessment of this proposal?

AAdopt it for performance-critical paths only; the coordinator remains the entry point but subagents can call siblings for known sub-workflows
BReject it; subagents cannot directly invoke other subagents in the standard pattern, and even if they could, hub-and-spoke is preferable for observability and error handlingCorrect
CAdopt it with a shared memory store so subagents can communicate without re-sending context through the coordinator
DReject it now, but plan to migrate to a peer-to-peer agent mesh once Anthropic releases A2A primitives

Why: In the standard Claude Agent SDK pattern, subagents are isolated — they can only be invoked from the parent that spawned them (typically the coordinator). Even setting that aside, the hub-and-spoke pattern is preferred for production: the coordinator sees every delegation and result, owns error handling and retries, and can rebalance work. A peer-to- peer mesh (A, D) loses observability. A shared memory store (C) introduces stale-read and consistency problems and doesn't solve the architectural concern.

A platform engineer is reviewing a multi-agent legal research system at scale. The architecture today:

A coordinator receives queries

For each query, the coordinator spawns four subagents (research, case-law, synthesis, report) in sequence

The coordinator forwards each subagent's full output as context to the next

Production telemetry shows:

p50 latency: 38 seconds

p99 latency: 92 seconds

~22% of queries are simple lookups ("What's the holding in Smith v. Jones?") that don't need synthesis or report-writing

~12% of queries are complex ("Compare the evolution of Fourth Amendment jurisprudence from 1960–2020 across three circuits") and genuinely benefit from the full pipeline

~66% are in between

The query mix is evolving as the product team adds new user segments

The platform engineer proposes four optimizations. Which is the most architecturally sound?

ABuild a query-complexity classifier (small fine-tuned model) that routes each query to one of three predefined subagent combinations (simple, medium, complex); retrain monthly as the query mix evolves
BHave the coordinator itself analyze each incoming query and dynamically decide which subagents to invoke and in what order — leveraging Claude's reasoning to adapt to whatever the query needs, including new query types that didn't exist at design timeCorrect
CRun all four subagents in parallel for every query, then have the synthesis subagent decide which outputs to use
DCache the output of each subagent keyed on a hash of its input prompt, so repeat queries skip subagents whose context hasn't changed

Why: The query mix is diverse, evolving, and includes new query types over time. This is the architecture pattern question: dynamic LLM-driven routing wins over predefined routing tables when the input distribution is non-stationary. A (classifier): A trained classifier requires labels, retraining, drift monitoring, and ops overhead — and it bakes in today's query categories. When the product team adds a new user segment that asks novel questions, the classifier silently mis- routes them. B (coordinator-driven dynamic routing): The coordinator LLM can reason about this specific query and pick just the subagents needed. New query types adapt naturally because reasoning is general, not categorical. C (run all four in parallel): Wastes 78% of the workload (everything below "complex") and burns tokens for no benefit. D (caching): Helps repeat queries but does nothing for the variable-complexity routing problem. Misses the architectural question. This is the strength of the orchestrator pattern — let the LLM coordinator make routing decisions rather than encoding them statically.

Build exercise

Build a research coordinator with gap-driven refinement

~55 min
  1. 1
    Write a coordinator that takes a broad query (e.g. "the creative industries") and decomposes it into distinct subtopic subagents with non-overlapping scope.

    Why: Partitioning by distinct subtopic/source type is what minimizes duplication and prevents coverage gaps.

  2. 2
    Have the coordinator dynamically decide how many subagents to spawn from the query's breadth, rather than always spawning a fixed set.

    You should see: A narrow query spawns fewer subagents; a broad one spawns more.

  3. 3
    Route every subagent result back through the coordinator and log it; never let a subagent call another subagent.

    Why: Central routing is what gives you observability and one place to handle errors.

  4. 4
    After the first synthesis, have the coordinator detect coverage gaps and re-delegate targeted subtasks for only the missing areas.

    You should see: The second pass fills gaps (e.g. performing arts, publishing) that the first decomposition missed.

  5. 5
    Deliberately narrow the initial decomposition to a single subtopic, then confirm the gap-detection loop — not a subagent change — is what restores full coverage.

    Why: This proves the root cause of incomplete coverage lives in the coordinator's decomposition.

Sources

Drill Agentic Architecture & Orchestration

Practice only this domain’s questions, untimed, with instant explanations.