Domain 4 · Task 4.5
Batch Processing Strategies
Design efficient batch processing strategies.
The Message Batches API processes requests asynchronously for 50% savings within a window of up to 24 hours, with no guaranteed latency SLA. That trade — half the cost for no latency guarantee — makes it ideal for non-blocking, latency-tolerant work and wrong for anything where a human or a pipeline is waiting on the result. The core skill is matching the API to the latency need and doing the SLA math, not batching everything because it is cheaper.
Key concept
Someone is waiting → synchronous. Results needed by morning or weekly → batch (50% cheaper, up to 24h, no SLA). Never batch a blocking gate like a pre-merge check.
What you need to know
Batch vs synchronous: the latency decision
The whole decision is: is anything blocked on this result?
| Message Batches API | Synchronous | |
|---|---|---|
| Cost | 50% savings | Full price |
| Latency | Up to 24h, no SLA | Immediate |
| Fits | Overnight reports, weekly audits, nightly test generation, bulk document processing | Pre-merge checks, interactive review — anything where someone waits |
Use batch for non-blocking, latency-tolerant work. Use synchronous for blocking workflows. A pre-merge check is the canonical thing you must never batch — a 24-hour window on a gate that blocks every merge is a non-starter regardless of the savings.
No multi-turn tool calling; correlate with custom_id
Two operational constraints define how you use batches:
- No multi-turn tool calling within a single batch request — you cannot run a tool mid-request and feed its result back for the model to continue. Each batch request is one-shot. Agentic loops that need tool round-trips stay synchronous.
custom_idlinks each request to its response. Because results come back asynchronously and possibly out of order, you tag every request with acustom_idand use it to correlate the response — and to resubmit only the failed items (for example, chunk a document that exceeded context and resubmit just that one).
batch = client.messages.batches.create(
requests=[
{
"custom_id": f"doc-{doc.id}", # correlate response; resubmit only failures
"params": {
"model": "claude-sonnet-4-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": doc.text}],
},
}
for doc in nightly_docs
]
)SLA math and sample-first refinement
When there is a deadline SLA, compute the submission frequency from the processing window. To guarantee a 30-hour SLA with a 24-hour processing window, submit within a 6-hour window before the deadline (and split into ~4-hour windows for frequent submissions so no item risks blowing the SLA). Before batching large volumes, refine your prompt on a sample first to maximize first-pass success — a systematic prompt error multiplied across thousands of batched requests wastes the whole run. On failures, identify them by custom_id and resubmit only those, rather than re-running the entire batch.
Exam traps
| The trap | The reality |
|---|---|
| The Batches API is 50% cheaper, so move both the overnight report and the pre-merge check to it. | Batch the overnight report only; keep the pre-merge check synchronous. A blocking gate can't tolerate a 24-hour window with no SLA. |
| Batches are usually fast enough, so they're fine for a blocking check with a timeout fallback. | "Often faster" and "timeout fallback" misunderstand the SLA — there is no latency guarantee. Never batch a blocking gate. |
| You can run multi-turn tool calls inside a single batch request. | Multi-turn tool calling is not supported within a batch request. Agentic loops that need tool round-trips must stay synchronous. |
| When a batch has failures, resubmit the whole batch. | Identify failures by custom_id and resubmit only those items (e.g. chunk an oversized document). Refine prompts on a sample before large runs. |
Practice scenario
Real questions from the bank that test this topic — the correct answer is highlighted.
You submit a batch of 5,000 extraction requests via the Message Batches API. After 6 hours, processing_status is "ended" . The results file contains:
- 4,800 with result.type: "succeeded"
- 150 with result.type: "errored" (various errors)
- 30 with result.type: "expired"
- 20 with result.type: "canceled"
What do these result types tell you, and what's the appropriate next action?
custom_id scontext_length_exceeded , validation errors); expired = the 24-hour processing window elapsed without the request being processed; canceled = the batch was canceled before those requests ran. Resubmit only the expired and errored requests (with fixes for the errors), keyed by their custom_id sCorrectWhy: Each request in a batch has its own result type. succeeded = normal completion. errored = per-request error (could be context_length_exceeded , invalid request, etc.). expired = the 24-hour batch SLO elapsed before the request was processed. canceled = the batch was canceled before this request ran. Every request has a custom_id so you can selectively resubmit specific failures with fixes. Wholesale resubmission (A, D) wastes the 4,800 successes and 50% discount. C invents a key-expiry semantics.
You're using the Message Batches API. After the batch ends, the result file shows mixed result types: succeeded , errored , expired , canceled . Specifically you see 30 expired records.
What does expired mean and what should you do with those 30?
custom_id s are preserved so you can resubmit just those 30 in a new batchCorrectWhy: Each batch request has its own result type. expired means that specific request's 24- hour processing window elapsed before it was processed — typically when the batch was very large or under heavy load. The custom_id of each expired request is preserved in the results, so you can resubmit just those records in a new batch. B (API key expiry) is invented semantics. C (model deprecation) is invented. D (batch-wide state) is wrong — expired is per-request.
Build exercise
Route a mixed workload between batch and synchronous
~40 min- 1List your two workloads: a pre-merge CI check (blocking) and an overnight tech-debt report (latency-tolerant).
Why: The routing decision is entirely about whether someone is waiting on the result.
- 2Submit the overnight report to the Message Batches API with a
custom_idper document.You should see: 50% cost savings and correlatable results, with completion within the 24-hour window.
- 3Keep the pre-merge check on the synchronous API and confirm it returns immediately.
Why: A blocking gate can't tolerate a 24-hour window with no SLA.
- 4Simulate a failure on one batch item and resubmit only that item by its
custom_id(e.g. after chunking an oversized document).Why: Resubmitting only failures avoids re-running the whole batch.
- 5Given a 30-hour deadline SLA and a 24-hour window, compute the latest safe submission time (a 6-hour buffer) and set your submission cadence.
You should see: A submission schedule that guarantees the SLA even at the worst-case 24-hour processing time.
Sources
Drill Prompt Engineering & Structured Output
Practice only this domain’s questions, untimed, with instant explanations.