CCAF logo

Domain 5 · Task 5.5

Human Review & Confidence Calibration

Design human-review workflows and confidence calibration.

"We hit 97% accuracy — can we drop the human review?" is the trap this lesson dismantles. A single headline number averages over document types and fields, and that average can hide a category where the system is 40% wrong. Designing a human-review workflow means measuring accuracy where it actually varies, calibrating confidence on labeled data, and continuing to sample even the outputs the system is sure about.

Key concept

Don't trust aggregate accuracy — segment by type and field; calibrate confidence on labeled data; sample high-confidence outputs continuously.

What you need to know

Aggregate accuracy is a mask

A 97% overall accuracy figure can conceal 40% error on one document type or one field, because the type is rare enough that its errors barely move the average. Before reducing human review, analyze accuracy by document type and by field segment — the invoice-total field on scanned PDFs may fail badly while the average looks fine. The decision to automate must rest on the worst-performing segment that matters, not the mean.

Field-level confidence, calibrated on labeled data

Route reviewer attention with field-level confidence scores — not one score per document, but per extracted field — and calibrate the thresholds on a labeled validation set so that "confidence 0.9" actually means ~90% correct. Calibrated confidence lets you send low-confidence or ambiguous/contradictory outputs to humans while auto-accepting the high-confidence ones, spending limited reviewer capacity where it changes outcomes. Note the contrast with Task 5.2: a model *self-rating* its confidence is unreliable, but confidence *calibrated against labeled ground truth* is a valid routing signal.

Stratified sampling to keep measuring

Calibration is not one-and-done — data drifts and novel document formats appear. Stratified random sampling of high-confidence extractions keeps measuring the true error rate in the outputs you're auto-accepting and detects novel patterns before they cause damage. Stratifying (sampling within each document type/field segment rather than uniformly) ensures rare-but-important categories are actually represented in the sample instead of being averaged away again.

Validate before you automate

The workflow order matters: validate by document type and field segment before automating, calibrate on labeled data, then automate only the segments that clear the bar — and keep sampling the rest. "Be conservative" or "only extract high-confidence values" as a prompt instruction is not a substitute; it doesn't produce a measurable, segmented error rate. The mechanism is empirical measurement and calibrated routing, not a softer instruction.

Exam traps

The trapThe reality
97% overall accuracy is well above our bar, so we can safely turn off human review.Aggregate accuracy can hide 40% error on a specific type or field. Segment accuracy by type and field before automating.
Confidence scores are unreliable for routing — Task 5.2 said so — so don't use them for review.Model self-rated confidence is unreliable, but field-level confidence calibrated on a labeled validation set is a valid routing signal. The difference is calibration.
Once the system is validated, you can stop reviewing high-confidence outputs entirely.Use stratified random sampling of high-confidence extractions to keep measuring error and catch novel patterns as data drifts.
Telling the model to "be conservative and only extract high-confidence fields" reduces the need for review.That's a vague instruction, not a measurement. You need per-type/per-field accuracy and calibrated confidence, not a softer prompt.

Practice scenario

Real questions from the bank that test this topic — the correct answer is highlighted.

Your team built an extraction system that achieves 96% accuracy overall on a labeled validation set. The team plans to automate high-confidence extractions (>90% confidence) routing only low-confidence ones to human review. Before deploying, your data analyst breaks down accuracy by document type:

Document TypeVolumeAccuracy
Standard business docs82%99.2%
Government documents11%95.1%
Hand-scanned receipts5%64.8%
Multilingual contracts2%41.2%

Your overall 96% comes from the weighted average. You're about to ship the automation.

AShip automation as planned; 96% overall accuracy is the contractually agreed threshold.
BApply uniform 90% confidence threshold across all document types as planned, but allocate extra reviewer capacity for the 60% of multilingual contracts that will fall below threshold.
CApply different confidence thresholds and review policies per segment: aggressive automation for business docs; conservative for government; require 100% human review for hand-scanned receipts and multilingual contracts regardless of confidence.Correct
DRun a 2-week pilot routing 25% of high-confidence extractions to automation; monitor for systematic errors and roll back if found.

Why: Aggregate accuracy hides segment failures. 64.8% on hand-scanned and 41.2% on multilingual would cause major downstream errors. Segment-level policies (C) are documented Task 5.5 practice. Uniform thresholds (A, B) ship known failures. Pilot (D) is reasonable but doesn't address the known segment problem.

Your system has been operating with 100% human review for 3 months. Analysis shows that extractions with model confidence >90% have 97% accuracy overall. To reduce reviewer workload, you plan to automate high-confidence extractions. Before deploying, what validation step is most critical?

AAnalyze accuracy by document type and field to verify high-confidence extractions perform consistently across all segments, not just in aggregate.Correct
BCompare accuracy at different confidence thresholds (85%, 90%, 95%) to find the optimal cutoff that maximizes automation while minimizing errors.
CRun a two-week pilot routing 25% of high-confidence extractions directly to downstream systems and monitor error reports.
DVerify that 97% accuracy meets requirements for all downstream systems that consume the extracted data.

Why: Aggregate accuracy hides segment failures — one document type could be 70% while others are 99%. Auto-routing by overall number alone risks systemic errors in the weak segments.

Build exercise

Design a segmented, sampled extraction-review workflow

~50 min
  1. 1
    Assemble a labeled validation set of extractions spanning several document types (e.g. invoices, receipts, statements) and several fields per type.

    Why: You can only measure per-segment accuracy and calibrate confidence against ground truth.

  2. 2
    Compute the aggregate accuracy, then break it down by document type and by field.

    You should see: A high overall number that hides at least one low-accuracy type or field — the masking effect made concrete.

  3. 3
    Have the extractor emit a field-level confidence score, and calibrate thresholds against the labeled set so the scores map to real accuracy.

    Why: Calibrated field-level confidence is the valid routing signal; an uncalibrated self-rating is not.

  4. 4
    Build a router that sends low-confidence, ambiguous, or contradictory fields to human review and auto-accepts the calibrated high-confidence ones.

    You should see: Reviewer effort concentrates on the fields where it changes the outcome.

  5. 5
    Add stratified random sampling of the auto-accepted high-confidence extractions, sampling within each type/field segment.

    Why: Stratifying ensures rare-but-important segments appear in the sample instead of being averaged away.

  6. 6
    Introduce a novel document format the system wasn't validated on and confirm the sampling surfaces its elevated error rate.

    You should see: The stratified sample flags the novel pattern before it silently corrupts the auto-accepted outputs.

Sources

Drill Context Management & Reliability

Practice only this domain’s questions, untimed, with instant explanations.