The Harness Era · 07 / 08

    The Trust Debt Index for Agents

    Enterprises need a way to quantify unsupported claims, weak evidence, and ungoverned automation.

    Published Jun 29, 2026·Essay · ~4 min·A KokoAI point of view

    Forrester gave the problem a name; the next move is to give it a number. The "trust tax" — the drag that ungoverned, unexplainable AI puts on adoption — is real, expensive, and so far unmeasured. Boards approve AI spend they cannot defend, risk committees veto deployments they cannot quantify, procurement waves through agents whose control posture nobody scored. A named tax that no one prices is a tax everyone overpays. The discipline that fixed the equivalent problem in counterparty risk was not better intentions. It was a score.

    Before the credit score, lending ran on relationships and judgment — illegible, unportable, impossible to compare across two borrowers. The score made counterparty risk a number: comparable, gate-able, tradable. A lender could decline below a threshold, price the spread above it, and sell the exposure to someone who read the same number the same way. Agentic automation is at the pre-score moment. Every enterprise is accumulating ungoverned automation it cannot compare, cannot price, and cannot gate. A Trust Debt Index does for agents what the credit score did for borrowers: it turns control posture into a number that production, procurement, and deal terms can all read.

    Trust debt is a balance, not a vibe

    Trust debt is the accumulated gap between what an agent asserts and what it can prove. It compounds like financial debt — every shortcut taken to ship faster is a liability carried into production, and the interest is paid in incidents, audit findings, and the deployments risk won't sign. The index makes the balance legible by scoring the signals that predict failure.

    SignalWhat it measures
    Unsupported claimsAssertions the agent makes with no source behind them
    Citation / evidence gapsOutputs that cite weakly, partially, or not at all
    Missing controlsNo segregation of duties, no approval gate on material actions
    Absent human-in-the-loopHigh-impact actions executed with no human accountable
    No evidence trailActions that leave no reconstructable record of what happened and why
    No eval pack / regression gateNo test suite, so quality drift ships silently
    Stale groundingDecisions made against context that has gone out of date

    Each signal is independently observable, and each maps to a failure mode a controller already understands. An agent that asserts without sourcing fails an audit; one with no approval gate fails segregation of duties; one with no eval pack degrades until a customer notices. Scored together, they produce one number per agent — and rolled up, one number per portfolio.

    The standard the index enforces

    A score is only useful against a standard. The index reads against a single governance bar, and every component above asks whether an agent clears it. An enterprise-ready agent is:

    • Always sourced — every claim traces to evidence, not to the model's confidence.
    • Always owned — a named human is accountable for what the agent does.
    • Always tested — an eval pack gates every change before it reaches production.
    • Always permissioned — the agent acts only within scoped, least-privilege authority.
    • Always observable — quality, cost, tool use, and exceptions are visible in real time.
    • Always value-linked — the agent carries a baseline, a target, and a lever it moves.
    • Human-accountable — material actions clear an approval gate before they execute.

    These are not aspirations. They are the seven conditions a low score certifies and a high one flags. An agent that meets all seven carries almost no trust debt; one that meets two carries a liability the index now states in a number instead of a reviewer's gut.

    A number that gates, prices, and travels

    A comparable score moves the decision out of the meeting and into the workflow. It does three things a named tax never could.

    It gates production. Set a threshold; agents above it deploy, agents below it go back to the harness. The CFO stops funding on faith, the risk committee stops vetoing on instinct — both read the same number. It prices procurement. The score becomes a buying term, the way a SOC 2 report already is; a vendor shipping high-trust-debt agents finds the tax is no longer abstract — it is a worse price or a lost deal. And it travels into deal terms. As agents touch revenue and cash, the score becomes diligence — an acquirer, an insurer, a partner all want the number before they sign.

    This is not hypothetical infrastructure. KokoKnows already runs a Trust Debt Index on public companies — turning control-posture signals disclosed in SEC filings into a comparable score across a universe of issuers. Extending that engine from corporate disclosure to individual agents is a change of input, not of method. The scoring discipline exists; agents are the next thing it reads.

    The harness is what the score reads from

    A Trust Debt Index is not a survey you fill out. It is a meter that reads instruments — and the harness installs them. Its components map one-to-one onto the artifacts the earlier pieces in this series built:

    • The agent card is where sourcing, ownership, and the value link are declared. No card, no provenance to read.
    • The control pack is where segregation of duties, approval gates, and permissioning live. No control pack, the missing-controls and human-in-the-loop signals fire by default.
    • The eval pack is the regression gate. No eval pack, the test signal fails.
    • The evidence ledger is the trail. No ledger, nothing to reconstruct, and the no-evidence-trail signal is automatic.

    Which makes the index a diagnosis, not just a grade. A high score is the direct, mechanical consequence of skipping the harness. The agent stood up as a clever script with no card, no controls, no tests, no ledger does not score badly because someone judged it harshly. It scores badly because there is nothing for the meter to read — and that absence is the liability the score exists to surface. Build the harness and the score takes care of itself. Skip it, and the index does not punish you; it shows you the bill you were already running up, signal by signal — which doubles as the remediation list.

    The trust tax was always going to get priced. The only question is whether the enterprise prices its own agents first, with a number it controls and can act on — or waits for a procurement team, an auditor, or an acquirer to price them instead, on terms it does not set.

    This essay is the seventh entry in the eight-part series, The Harness Era. The full-length analysis — the component scoring model, the credit-score analogy in full, and how the score reads from the harness — is linked below. The finale brings the series home: Agentic Finance Is the First Killer App — why the office of the CFO is where governed agentic execution proves itself first.

    Go deeper

    Read the full-length analysis

    The market structure, the architecture, and the evidence behind the thesis — the source this essay draws on.

    Open the full analysis →