Forrester named the cost and stopped there. The "trust tax" — the drag that ungoverned, unexplainable AI puts on adoption — is now a fixture of the analyst vocabulary and the boardroom conversation. It is also still a metaphor. Everyone agrees the tax is real and that it is slowing deployment, but no one can tell you what a given agent's trust tax actually is, in a number two people would compute the same way. That gap is not cosmetic. A named tax that cannot be priced is a tax that gets paid in the dark: budgets approved on faith, deployments vetoed on instinct, agents waved through procurement on a demo. This analysis makes the case that the next step is the obvious one — give the trust tax a number — and lays out what that number measures, what standard it enforces, where it gates, and why it is a direct readout of the harness.
1. The pre-score moment
Counterparty risk used to be a matter of judgment. A lender knew a borrower, or knew someone who did, and made a call. The call was illegible to anyone outside the relationship, unportable from one institution to the next, and impossible to compare across two applicants without re-doing the judgment from scratch. Credit was extended on confidence and withdrawn on rumor, and the whole system priced risk badly because it had no common unit to price it in.
The credit score changed the kind of thing counterparty risk was. It took a fuzzy, relationship-bound judgment and made it a number — one that meant the same thing to every lender, that a borrower carried from institution to institution, and that could be compared, thresholded, and traded. A lender could decline below a cutoff, price a spread above it, and sell the exposure to a third party who read the identical number and underwrote it the identical way. The score did not make borrowers more trustworthy. It made trustworthiness legible, and legibility is what let an entire market form on top of it.
Agentic automation is at the pre-score moment. Every enterprise is accumulating agents, copilots, scripts, and workflows whose control posture nobody has scored. The trustworthiness of any one of them is a matter of whoever built it, knowable only by re-doing the review, and impossible to compare against the agent the next team stood up. The enterprise has automation it cannot compare, cannot price, and cannot gate — exactly the condition lending was in before the score. A Trust Debt Index is the score that ends that condition for agents.
2. Trust debt, defined
Trust debt is the accumulated gap between what an agent asserts and what it can prove. The word debt is precise, not decorative. Like financial debt, it is incurred to move faster — every control skipped, every test not written, every source not captured is a shortcut that ships the agent sooner. Like financial debt, it is a liability carried forward, not a one-time cost. And like financial debt, it accrues interest: the bill arrives later, as incidents in production, as findings in the audit, as the deployment the risk committee will not approve, as the deal term a buyer marks down. An enterprise running a fleet of ungoverned agents is not operating cheaply. It is operating on credit, and the principal is coming due.
The reason to score it rather than describe it is the same reason credit moved from judgment to number. A description ("this agent is risky") cannot gate, cannot price, and cannot be compared. A score can do all three. And because trust debt is built from independently observable signals, it can be computed mechanically rather than adjudicated case by case — which is the only way it scales past the handful of agents a human reviewer can hold in their head.
3. The components of trust debt
The index reads seven signals. Each is independently observable, each maps to a failure mode a controller already recognizes, and each corresponds to something the harness either produces or, by its absence, leaves missing.
| Signal | What it measures | Failure mode when present |
|---|---|---|
| Unsupported claims | Assertions the agent makes with no source behind them | Confident output that cannot survive a "how do you know?" — the audit fails on the first sampled decision |
| Citation / evidence gaps | Outputs that cite weakly, partially, or after the fact | The reasoning cannot be reconstructed; reviewers must re-derive every conclusion by hand |
| Missing controls | No segregation of duties, no approval gate on material actions | One actor initiates and approves; the classic control breakdown, now automated and faster |
| Absent human-in-the-loop | Material, high-impact actions executed with no human accountable | An irreversible action ships with no one who can be asked why, or stop it |
| No evidence trail | Actions that leave no reconstructable record of inputs, reasoning, and result | Nothing to audit, nothing to replay after an incident, no defense in a dispute |
| No eval pack / regression gate | No test suite gating changes | Quality drifts silently; the agent degrades and the first to notice is a customer or a regulator |
| Stale grounding | Decisions made against context that has gone out of date | The agent is confidently, fluently wrong because the world moved and its grounding did not |
These are not weighted guesses. Each is a yes/no or a measurable degree, observable from the agent's own artifacts and runtime telemetry. An agent that sources every claim, cites cleanly, runs under an approval gate, keeps a human accountable for material actions, writes a complete evidence trail, passes an eval pack on every change, and grounds against fresh context carries almost no trust debt. An agent missing four of the seven carries a liability the index now states as a number, instead of leaving it to whether a reviewer happened to ask the right question.
The signals are also additive and rolled-up. One agent gets a score. A process domain — lead-to-cash, source-to-pay, record-to-report — gets the aggregate of its agents' scores. A portfolio gets the aggregate across domains. The same number works at every altitude, which is what lets a CFO talk about portfolio trust debt and a builder talk about a single agent's, in the same unit.
4. The standard the index enforces
A score has to measure against something. The Trust Debt Index reads against a single governance standard, and every component in section 3 is one way of asking whether the agent clears it. An enterprise-ready agent is:
- Always sourced. Every claim traces to evidence the enterprise can inspect — a document, a record, a system of truth — not to the model's fluency. Sourcing is the difference between a decision and an opinion.
- Always owned. A named human owns the agent's behavior and outcomes. Ownership is what turns "the AI did it" back into "someone is accountable for what it did."
- Always tested. An eval pack gates every change. The agent that worked yesterday is re-proven today, before the change reaches anyone who depends on it.
- Always permissioned. The agent operates inside scoped, least-privilege authority. It can do what its job requires and provably nothing more.
- Always observable. Quality, cost, tool use, latency, and policy exceptions are visible in real time, not reconstructed after an incident.
- Always value-linked. The agent carries a baseline, a target, and one of the five value levers it moves — revenue, margin, working capital, productivity, or risk. An agent with no value link is an expense; the standard requires it be an investment.
- Human-accountable. Material actions clear an approval gate before they execute. Automation does not mean unattended; it means the human attention is spent where it changes the outcome.
These seven conditions are the bar. A low trust debt score certifies that an agent meets them. A high score flags exactly which it fails. The standard is what makes the number mean the same thing across two agents built by two teams who never spoke — the property that made the credit score useful and the property the enterprise's current case-by-case reviews can never have.
5. A number that gates, prices, and travels
The value of a comparable score is that it relocates the decision. A trust judgment made in a meeting is slow, inconsistent, and gone the moment the meeting ends. A trust number lives in the workflow and does three things the named tax never could.
It gates production. The enterprise sets a threshold. Agents above it deploy; agents below it go back to the harness with a list of exactly which signals to fix. The funding conversation changes shape: the CFO is no longer asked to approve "AI" on faith, but to approve agents that cleared a stated bar, with the trust debt of the portfolio reported alongside its value. The risk committee stops being the office of "no" and becomes the office of "here is the threshold" — because for the first time it has a number to set the threshold against. Production access becomes a function of score, the way a deploy gate is a function of a passing test suite.
It prices procurement. When an enterprise buys an agent or an agentic platform, the supplier's trust debt score becomes a procurement term, sitting next to the security questionnaire and the SOC 2 report that buyers already demand. A vendor shipping high-trust-debt agents — unsourced, untested, uncontrolled — finds the trust tax is no longer an abstraction in an analyst report. It is a worse price, a longer security review, or a lost deal. The market pressure that the credit score put on borrowers, the trust debt index puts on agent vendors: clear the bar or pay the spread.
It travels into deal terms. As agents move from drafting emails to touching revenue, cash, and the close, their trust debt becomes diligence. An acquirer valuing a target now has an automation estate to assess, and a portfolio trust debt score is the fastest way to assess it — a high score is a remediation cost priced into the offer. An insurer underwriting AI-related coverage wants the number. A partner integrating its systems with the enterprise's agents wants the number. The score becomes the unit in which agentic risk is bought, sold, and insured — the trust tax, finally priced, the way counterparty risk became priceable once it had a score.
6. Prior art: the index already exists
This is not a thought experiment about infrastructure that would have to be invented. KokoKnows already runs a Trust Debt Index on public companies. It turns control-posture signals disclosed in SEC filings — the language and substance of how a company governs itself — into a comparable score across an entire universe of issuers. The hard parts are already solved: defining signals that predict failure, extracting them at scale from messy real-world inputs, normalizing them into a number that means the same thing across very different entities, and publishing it as something comparable.
Extending that engine from corporate disclosure to individual agents is a change of input, not of method. Where the corporate index reads filings, the agent index reads the harness — agent cards, control packs, eval packs, evidence ledgers, observability telemetry. The scoring discipline, the normalization, the comparability are the same machinery pointed at a new, richer, and far more structured source of signal. An agent's control posture is not buried in prose the way a company's is; it is declared in artifacts built for exactly this purpose. The agent index is, if anything, the easier read.
7. The harness is what the score reads from
A Trust Debt Index is not a questionnaire. It is a meter, and a meter is only as good as the instruments it reads. The harness is what installs those instruments, which is why the components of the index map one-to-one onto the artifacts the earlier pieces of this series built:
- The agent card declares and carries the agent's sourcing, ownership, scope, and value link. It is where the index reads provenance, the named owner, and the value lever. No card means no provenance to read — the unsupported-claims and value-link signals fail by default.
- The control pack holds segregation of duties, approval gates, and least-privilege permissioning. It is where the index reads the controls and human-in-the-loop posture. No control pack means the missing-controls and absent-human signals fire automatically, because there is nothing asserting otherwise.
- The eval pack is the regression gate. It is where the index reads whether changes are tested. No eval pack means the test signal fails on contact — there is no suite for the meter to confirm.
- The evidence ledger is the trail. It is where the index reads reconstructability. No ledger means the no-evidence-trail signal is automatic, because there is, definitionally, nothing to reconstruct.
This mapping is the analysis's central claim, and it has a sharp consequence: a high trust debt score is the direct, mechanical result of skipping the harness. The agent stood up as a clever script — no card, no control pack, no eval pack, no ledger — does not score badly because a reviewer was in a bad mood. It scores badly because there is nothing for the meter to read, and that absence is the liability the score exists to surface. The index does not add a judgment on top of the agent. It reads what the agent's own construction left present or missing.
Which reframes the index from grade to diagnosis. A high score does not just say "this agent is risky." It says, signal by signal, which piece of the harness is missing and therefore which artifact to build to bring the score down. The remediation list writes itself: low on the evidence signal, install the ledger; low on the test signal, write the eval pack; low on controls, add the control pack and the approval gate. Build the harness and the score takes care of itself, because the score was only ever reading the harness. Skip it, and the index does not punish the enterprise. It just states the bill the enterprise was already running up, in a number it can finally act on.
8. The forward view
The trust tax was always going to be priced. The only open question was who would price the enterprise's agents first — and on whose terms. Through the next several quarters, expect trust debt to move from metaphor to metric: a number reported alongside agent value, a gate on production access, a line in the procurement template. The enterprises that adopt it early will do so because it is the only way to say yes to agentic automation at scale without losing the thread of what those agents are actually allowed to do and prove.
Further out, the score does what the credit score did — it leaves the building. Procurement teams demand it from vendors. Auditors expect it as standard evidence. Acquirers price it into diligence. Insurers underwrite against it. At that point the enterprise that scored its own agents first, with an index it controls and can act on, is in the position the prepared borrower is in: it walks into the conversation with the number already in hand, on terms it set, while its competitors wait for someone less kind to assign them one. The harness produces the agents worth a high score. The Trust Debt Index is how the rest of the world finally learns to read it.