CCAF logo

Domain 1 · Task 1.4

Workflow Enforcement & Handoff

Implement multi-step workflows with enforcement and handoff patterns.

Multi-step workflows raise one central question: how do you *guarantee* a required step happens before a later one? The answer separates programmatic enforcement (hooks, prerequisite gates) from prompt-based guidance. When deterministic compliance is required — for example verifying identity before any financial operation — prompt instructions carry a non-zero failure rate, and that failure has real cost.

Key concept

Mandatory ordering with financial or compliance stakes → programmatic gate, not a prompt. "Would this fail occasionally even if a careful human wrote the perfect instruction?" If yes and it's costly, enforce it structurally.

What you need to know

Programmatic enforcement vs prompt guidance

Prompt-based guidance — even a strongly worded system prompt or few-shot examples — is probabilistic: it works most of the time, not every time. Programmatic enforcement is deterministic: a PreToolUse hook or prerequisite gate blocks a downstream tool call until its prerequisite has completed. The discrimination the exam rewards: if the behaviour *fails measurably despite instruction* and the failure has financial, security, or compliance consequences, you need a gate — not better wording.

Prerequisite gates

A prerequisite gate blocks specific tool calls until a precondition is satisfied. The canonical example: block process_refund (and lookup_order) until get_customer has returned a verified customer_id. Because the gate runs in code on every call, the ordering is guaranteed rather than merely encouraged.

python
@hook("PreToolUse")
def require_verified_customer(call, state):
    if call.name in ("lookup_order", "process_refund") and not state.verified_customer_id:
        return block(call, reason="Call get_customer first to verify identity.")

Decompose multi-concern requests

When one request bundles several concerns, decompose it into distinct items, investigate each — in parallel where possible, using shared context — then synthesize one unified resolution. This keeps each concern from being dropped and lets you handle them concurrently without losing the thread that ties them together.

Structured handoff protocols

Mid-process escalation to a human needs a structured handoff summary, because the human agent cannot see the transcript. Compile the concrete facts they need to act: customer ID, root-cause analysis, refund amount, and a recommended action. A handoff that just says "escalating this issue" forces the human to reconstruct everything from scratch — the summary must carry the details.

Exam traps

The trapThe reality
Strengthen the system prompt to say "verification is mandatory" to stop the agent skipping it.A prompt is probabilistic and will still fail some fraction of the time. With financial stakes, use a programmatic prerequisite gate.
Adding few-shot examples of the correct order reliably enforces a mandatory sequence.Few-shot is still probabilistic. Mandatory ordering with compliance stakes requires a hook or gate, not examples.
A routing classifier that sends refunds to the right handler fixes the skipped-verification bug.A classifier solves *availability/routing*, not *ordering*. The gate is what guarantees verification runs first.
Escalating to a human just means flagging the conversation for review.The human can't see the transcript. Compile a structured summary: customer ID, root cause, refund amount, recommended action.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your customer support agent must verify a customer's identity before processing any refund. Despite a system prompt instruction stating "Always call get_customer and verify the returned customer ID before calling process_refund ," production logs show that in 1.5% of cases the agent calls process_refund without first calling get_customer . Each occurrence corresponds to a financial error.

What is the most appropriate change?

AAdd a few-shot example to the system prompt showing the correct ordering, with explicit reasoning text demonstrating why verification must happen first.
BImplement a PreToolUse hook on process_refund that blocks the call and returns an error to the agent unless get_customer has returned a verified customer ID earlier in the conversation.Correct
CModify the process_refund MCP tool to internally call get_customer first and verify identity, returning an isError: true result if verification was not previously performed.
DIncrease the priority of the verification instruction by placing it at the very top of the system prompt and using emphatic language ("CRITICAL: Never skip verification").

Why: The guide is explicit: when deterministic compliance is required (e.g., identity verification before financial operations), prompt instructions alone have a non-zero failure rate. A PreToolUse hook that programmatically blocks process_refund until get_customer has succeeded provides the deterministic guarantee. Option A (few-shot) and Option D (emphatic prompt) are both probabilistic. Option C is tempting — moving the check into the tool itself — but it couples authentication logic into the refund tool and still depends on conversation state the tool can't reliably see; the canonical pattern (Task Statement 1.4) is programmatic prerequisites via hooks.

You're handling escalation logic for escalate_to_human . The human agent who picks up the case does not have access to the conversation transcript or tool results. After an extensive 30-turn investigation, the agent determines a $920 refund is owed due to a duplicate-charge bug in your payment gateway.

APass the full conversation history (messages array) so the human agent has complete context.
BPass the customer's original verbatim complaint plus the raw tool result objects from lookup_order and get_customer .
CPass a structured handoff summary including customer ID, root cause diagnosis (duplicate charge from gateway timeout), refund amount, and recommended action (issue refund + flag gateway logs for engineering).Correct
DPass a reference ID pointing to the persisted transcript in your case management database, so the human can fetch the full context if needed.

Why: Task Statement 1.4 — structured handoff including customer ID, root cause, refund amount, recommended action. Human pickup is faster, full transcript dumps force re- investigation.

Build exercise

Enforce identity verification with a prerequisite gate

~50 min
  1. 1
    Build a support agent with get_customer, lookup_order, and process_refund tools, and a system prompt that *asks* it to verify identity first.

    Why: You need the prompt-only baseline to observe that it skips verification some of the time.

  2. 2
    Run it across many simulated conversations and measure how often process_refund fires without a verified customer_id.

    You should see: A non-zero skip rate — the probabilistic failure the exam's canonical question describes (~12%).

  3. 3
    Add a PreToolUse hook that blocks lookup_order and process_refund until state carries a verified customer_id from get_customer.

    Why: The gate makes the ordering deterministic instead of relying on the model complying.

  4. 4
    Re-run the same conversations and confirm the skip rate drops to zero.

    You should see: Every refund is now preceded by verification — a hard guarantee.

  5. 5
    On escalation, have the agent emit a structured handoff summary (customer ID, root cause, refund amount, recommended action) rather than a bare "please review".

    Why: The human agent acts on the summary, not the transcript they can't see.

Sources

Drill Agentic Architecture & Orchestration

Practice only this domain’s questions, untimed, with instant explanations.