Friday, August 21, 2026

    Design human review workflows and confidence calibration

    Domain 5 — Context Management & Reliability · 15% of the exam

    Task statement 5.5

    Design human review workflows and confidence calibration

    What you should be able to do
    • Explain how an aggregate accuracy figure hides poor performance on one document type or field.
    • Use stratified random sampling of high-confidence output to measure real error rates.
    • Calibrate field-level confidence against a labelled validation set before trusting it for routing.
    • Route low-confidence and internally contradictory sources to human review first.
    Exam traps (4)

    Overall accuracy is high, so human review can be reduced.

    An aggregate averages over segments, so one document type failing badly is invisible behind the volume of easy ones — and that is where the review was still needed.

    Verify accuracy by document type and by field before reducing review anywhere.

    Review the low-confidence extractions and let the high-confidence ones through.

    Nothing then measures the high-confidence population, so a new error pattern inside it is never observed.

    Sample the high-confidence stream continuously — that is what detects novel errors.

    Use the model's confidence score as the review threshold.

    Uncalibrated confidence has no fixed relationship to accuracy, so the threshold means something different per field and per document type.

    Calibrate against a labelled validation set, then set thresholds from that.

    Review a random sample so the measurement is unbiased.

    A uniform sample is dominated by the common easy cases, so the rare segments that carry the risk are barely represented.

    Stratify the sample so each segment is measured, not just the bulk.

    Primary sources
    Know cold
    • Confidence calibration
    • Stratified sampling
    • Validator
    • Escalation

    7 practice questions in the bank are tagged to this task statement.

    Practise this task in context: open it inside the interactive study guide, which carries the concept cards, the mock quiz and the practice simulation.

    Blueprint-aligned independent practice. Koko's scenarios and company facts are fictional and synthetic. This aid does not reproduce official exam questions or Anthropic's undisclosed scoring model, and is not affiliated with or endorsed by Anthropic.