Monday, August 24, 2026

    Evaluation, Testing & Optimization

    16% of the exam · 6 published objectives

    Domain 4 · Evaluation, Testing & Optimization

    16%

    Weight as published in the exam guide. Your score report shows a percentage per domain and no domain has its own pass line — only the total scaled score decides the result.

    6 objectives, as published

    The guide lists these as bullets and assigns them no numbers. Nothing here numbers them either, and no score report will.

    • Define evaluation metrics (accuracy, latency, cost, safety, security)
    • Design evaluation datasets and test frameworks using mixed methodologies
    • Conduct A/B testing and iterative improvements
    • Diagnose system issues (prompt failure, hallucinations, model mismatch)
    • Optimize token usage, latency, and cost-performance trade-offs
    • Monitor system performance using logging and observability tools

    5 concepts — tap a card for the example, the right instinct and the trap

    Objectives in this domain Anthropic does not document

    • Observability — one page, and it is on the Claude Code host
      How we know: Observability with OpenTelemetry, under the Agent SDK, is the only observability page in the corpus. There is no platform-level equivalent, and the exam asks about monitoring strategies at scale.
      Read instead: That page for the mechanism, plus the Usage & Cost and Analytics APIs for what is measurable about a deployment. Nothing first-party discusses observability design for a multi-agent system.
      Objectives this leaves unsupported
      • Analyze observability challenges and select monitoring strategies at scale
      • Monitor system performance using logging and observability tools

    Working the whole blueprint? The study guide carries the eligibility gate, the exam logistics and every concept card in one place.