Domain 4 · Evaluation, Testing & Optimization
16%Weight as published in the exam guide. Your score report shows a percentage per domain and no domain has its own pass line — only the total scaled score decides the result.
6 objectives, as published
The guide lists these as bullets and assigns them no numbers. Nothing here numbers them either, and no score report will.
- Define evaluation metrics (accuracy, latency, cost, safety, security)
- Design evaluation datasets and test frameworks using mixed methodologies
- Conduct A/B testing and iterative improvements
- Diagnose system issues (prompt failure, hallucinations, model mismatch)
- Optimize token usage, latency, and cost-performance trade-offs
- Monitor system performance using logging and observability tools
5 concepts — tap a card for the example, the right instinct and the trap
Verified first-party sources for this domain
10 pages, each fetched and observed to resolve on 2026-08-24.
Objectives in this domain Anthropic does not document
- Observability — one page, and it is on the Claude Code hostHow we know: Observability with OpenTelemetry, under the Agent SDK, is the only observability page in the corpus. There is no platform-level equivalent, and the exam asks about monitoring strategies at scale.Read instead: That page for the mechanism, plus the Usage & Cost and Analytics APIs for what is measurable about a deployment. Nothing first-party discusses observability design for a multi-agent system.Objectives this leaves unsupported
- Analyze observability challenges and select monitoring strategies at scale
- Monitor system performance using logging and observability tools
Working the whole blueprint? The study guide carries the eligibility gate, the exam logistics and every concept card in one place.