Monday, August 31, 2026

    How the Koko Practice Index is calculated

    The formula, what it cannot tell you, and the answer-cue numbers for every pool this site serves

    All Claude certification study guides

    The Koko Practice Index is a linear study gauge. It is not an official exam score, it does not reproduce Anthropic's undisclosed scaling, and it is not a prediction.

    What it measures

    One thing: the share of practice items you answered correctly. That share is re-projected onto a 100–1000 range by a straight line, with no other input.

    index = round(100 + (percent correct ÷ 100) × 900)

    Difficulty does not enter it. Neither does domain weight, how discriminating an item turned out to be, how long you took, or how many attempts you have had. Two learners with the same percentage get the same index, whatever they drew.

    Percent correct mapped onto the Koko Practice Index
    Percent correctIndexWhat that row is
    0%100Nothing correct — the floor of the scale, which is not zero.
    25%325About what guessing returns on a four-option single-select item.
    50%550Half the items.
    75%775Three items in four.
    100%1000Every item correct.

    Where 720 comes from

    720 is Anthropic's published pass mark, set by a formal standard-setting study against Anthropic's own items. Koko reuses the number as a Koko study target and nothing more: it marks the point on this practice gauge that is worth aiming past before you book a seat. Reaching it here is evidence about this question bank.

    All 3 certifications covered here publish the same 100–1000 scale and the same 720 cut, which is why one index is applied to all of them. On this linear map the target is first reached at 69% correct, which is a different number of items on each exam:

    ExamItemsCorrect to reach 720Index earned
    CCDV-F5337728
    CCAR-F6042730
    CCAR-P6344729

    What it does not tell you

    The index borrows the official instrument's coordinates, so it looks like a scaled score wherever it appears. It is not one, and these are the specific claims it does not support.

    • It is not an Anthropic scaled score.

      Anthropic does not publish how performance maps onto its 100-1000 scale. Koko cannot reproduce an undisclosed formula, so it does not try — it uses a linear map and says so.

    • It is not calibrated against real exam outcomes.

      Calibration would need pairs of Koko results and official results from the same people. Koko holds none, so no statement about how an index maps onto a real sitting is available — including this one being close.

    • It does not predict whether you will pass.

      The published 720 cut is criterion-referenced from a formal standard-setting study on Anthropic's items. Reaching 720 here means you answered enough of Koko's practice items; it is evidence about this bank, not about that exam.

    • Two sittings at the same index are not equally hard.

      The index is percent correct and nothing else. Difficulty, domain weight and how discriminating an item is do not enter it, so an easy draw and a hard draw at the same percentage read identically.

    • It is not a blind measurement of you.

      Practice items are drawn from a fixed pool you can re-sit, review with rationales, and pause. That is what makes it useful for study and what makes it a poor estimate of a one-shot proctored result.

    • It is not scored on the exam's own items.

      Koko's items are original, written from the published blueprint and public documentation. No live, recalled or third-party exam content is used, so nothing here measures performance on the questions Anthropic actually asks.

    The answer-cue controls

    A practice score is worthless if the items can be guessed from their shape. This site's own bank once failed that badly — the correct option was the single longest in 96% of Mock items, and “always pick A” scored 83% on its legacy questions. Both were measured, both were fixed, and the linter that caught them now runs over every pool. These are its rules, with the live bounds:

    ControlThe cue it removesThe rule
    CUE-001The longest answer is usually the right one.A keyed option may not exceed 1.25x the item's median option WORD count unless some distractor reaches 85% of its length.
    CUE-002The same cue, measured the way a taker actually sees it — on screen length.The same rule on CHARACTERS at 1.4x / 75%, plus a hard 240-character ceiling on any option.
    CUE-003"All of the above" collapses four judgements into one.No option may use an of-the-above construction.
    CUE-004Absolutes mark the distractors, so the measured option stands out.always / never / only / completely / guaranteed may not appear in the key alone.
    CUE-005A distractor nobody bothered to make plausible is not a distractor.Every option carries its own rationale, and every distractor names the misconception it catches.
    CUE-007Guessing how many to select is not the skill being tested.A multiple-response item states its own count ("Select TWO." or "Select 2.").
    REF-001A rationale saying "B would be correct if…" cites an option that was relabeled at draw time.No rationale may name an option letter. Options are shuffled and relabeled when the item is drawn.
    BANK-POS"Always pick A" once scored 83% on this repo's own legacy items.No single answer position may hold more than 40% of a pool's single-select keys.
    BANK-LENKeys that are systematically longer than distractors, pool-wide.Mean keyed vs mean distractor length may differ by at most 12%.

    What those rules measure today

    Recomputed every time this page renders, over the same question objects the drills serve. cue-w is the share of items whose key is the single longest option by word count — four-option chance is 25%. rank is the keyed option's mean length position, where 0.5 is no advantage and 1.0 means the key is longest every time. diff is how far mean keyed length sits from mean distractor length, and pos is the largest share of keys held by one answer position.

    PoolItemscue-wcue-crankdiff-wdiff-cpos
    CCAR-F Mock exam and domain quizzes3859%42%0.6520%5%26%
    CCDV-F Practice drills10623%63%0.8076%10%26%
    CCAR-P Practice drills12631%65%0.8379%11%27%
    official CCAR-F sample items (reference)1217%58%0.7505%10%

    The last row is the calibration set: the 12 sample items printed in the official CCAR-F exam guide, run through the same functions. Every gated bound was checked to clear it first — a gate the real exam would fail is a gate somebody switches off. That is also why cue-c is published and never gated: the real exam measures 58% on characters, so a 40% character gate would reject Anthropic's own questions. 12 items is a small sample, and it is the largest published one available.

    CCAR-F mock exam and domain quizzes: within every gated bound. Some figures still sit above the official sample — the bounds are ceilings, not parity targets, and the table above is the honest comparison.

    CCDV-F practice drills: within every gated bound. Some figures still sit above the official sample — the bounds are ceilings, not parity targets, and the table above is the honest comparison.

    CCAR-P practice drills: within every gated bound. Some figures still sit above the official sample — the bounds are ceilings, not parity targets, and the table above is the honest comparison.

    What this page cannot measure

    The table above covers the pools your browser is holding — the drills and the Mock exam. It does not cover the simulated exam. Those items are graded on the server precisely so their answer keys never reach a browser, and a page that could measure them would be a page that had published them. They are measured instead where the keys live: by the build gate that refuses to regenerate the bank when a cue exceeds its bound, and by scripts/quiz-bias-report.ts, which prints every pool including the server-only ones. Both run the functions used above.

    Where the index appears

    On the simulated-exam result screen, in your saved results, and in the report you can download. It is the same number in all three, computed the same way. Back to the study guides.

    Powered by KokoAI