First time here? Four cards, under two minutes — then it's all yours.
What this page is
A working governance suite for AI that acts — agents that plan, call tools and change systems. Everything runs in your browser against a fixed control framework: nothing is uploaded, nothing needs an account, and the same answers always produce the same result.
What you walk away with
Three concrete artifacts, plus the reference catalogs behind them:
- A 22-domain maturity profile with your weakest domains and their accountable owners
- An A0–A5 agency-risk tier for any agent or use case — with the oversight model and control profile that tier demands
- A board-forwardable governance brief (.md or print/PDF) and a pre-filled DRAFT agent system card
Send it to a colleague
Every assessment has a Share button — text or email a link and the recipient's page recomputes your inputs through the same engine. Nothing is stored server-side; the link itself carries the answers.
Where it connects
Type a company once and it follows you across Koko: the classifier hands off to AI Stack Economics to price the control cost of your tier, the vendor landscape links back to stack pricing, and the Agent Registry view shows Koko's own production fleet governed by this exact model.
Agentic governance is the system of decision rights, policies, accountability, lifecycle gates, controls and assurance for AI that can plan, use tools, delegate and act. A control plane is what turns that intent into something machines enforce: discoverable agents, hard policy, runtime decisions, human approvals, telemetry, evidence and intervention.
The market is filling fast with platform-specific control planes — Microsoft Agent 365, Amazon Bedrock AgentCore, Google's Gemini Enterprise agent platform, ServiceNow AI Control Tower, IBM watsonx.governance, Salesforce Agentforce. Each governs its own runtime well and nobody else's. The durable position is the layer those platforms can't be: a vendor-neutral trust, control and value plane — one graph of agents, identities, controls, evidence, cost and outcomes; regulation converted into a reusable control library; risk assessed at design time and per action; and governance wired into the economics, so the enterprise knows not just is this agent safe but is it worth scaling.
Everything on this page runs on that thesis, in your browser, with no sign-up: a maturity diagnostic, a deterministic risk classifier, the full control catalog and regulatory crosswalk, the KPI dictionary — and Koko's own agent fleet as the worked example of what a governed registry actually looks like.
Your Governance Studio
The studio layer, working now — no sign-up, all in your browser. Assess, classify, and it generates your governance pack. Three passes:
- Pass 1
Score your maturity
22 governance domains, 0–5, against evidence you can defend.
- Pass 2
Classify an agent's risk
11 deterministic factors → an A0–A5 agency tier and control profile.
- Pass 3
Generate your pack
Maturity, agency verdict, control profile and a pre-filled system card — board-forwardable, in one file.
Complete 1 or 2
This is the studio layer — assessment, classification and pack generation, live and free. The connected estate graph, evidence hub and policy federation are the backend layers it grows into.
The model in one view
Five planes. Intent flows down; evidence flows back up into management and the next intent decision — governance as a loop, not a binder.
Ten design principles
Govern actions, not only outputs
Controls attach to plans, delegations, tool calls, transactions and side effects — not just to generated text.
Every agent is a principal
A unique identity, owner, purpose, version, risk tier, authority and expiry — like any employee or service account.
Authorization is contextual and continuous
A deployment approval is not a standing right to perform every future action.
Deterministic controls bound probabilistic systems
Material, irreversible or regulated actions use hard policy, never model judgment alone.
Least agency joins least privilege
Limit goals, tools, steps, sub-agents, spend, time, data, destinations and transaction values.
Human oversight is risk-triggered
The agent never decides whether its own action requires human approval.
Evidence is a product output
Every consequential action must be reconstructable without relying on hidden chain-of-thought.
Control cost is part of AI economics
A use case does not scale until value remains attractive after control, review, failure and change costs.
Federate before replacing
Reuse enterprise IAM, GRC, CMDB, SIEM, data governance, DevSecOps and native platform controls.
Open standards, portable records
Support MCP, A2A, OpenTelemetry and exportable policies — without assuming protocols alone create trust.
What's in the suite
Templates & artifacts
18 downloadable .mdThe working documents of the framework — system cards, risk assessments, oversight matrices, policies, playbooks — plus four completed synthetic samples so you can see the finished shape before you start. Take them; they're the point.
Agent System Card / Passport
templateThe authoritative per-agent record: identity, purpose, agency, actions, owners, limits.
Preview
--- template: agent_system_card version: 1.0 status: draft --- # Agent System Card / Passport ## Identification | Field | Value | |---|---| | AI system ID and name | | | Agent ID, name and version | | | Registry status / environment | | | Release hash and date | | | Parent / child agents | | | Sponsor / business owner | | | Product / technical owner | | | Data / risk / control owners | | | Vendor(s) and runtime | | ## Purpose and business context - Intended purpose: - Business process, subprocess and task: - In-scope users and affected stakeholders: - Business outcome and KPI baseline: - Allowed uses: - Prohibited uses: - Known limitations and required user disclosures: ## Agency and actions | Field | Value | |---|---| | Static agency tier | | | Autonomy mode | Assist / Recommend / Draft / Bounded action / Conditional autonomy | | Triggers and maximum duration | | | Maximum steps / retries / sub-agents | | | Permitted tools and actions | | | Financial / volume / rate limits | | | Irreversible or external actions | | | Required human approval | | | Segregation-of-duties design | | | Reversal / compensating action | | ## Identity, delegation and access - Agent/workload identity: - Requesting user or service context: - Acting-on-behalf-of chain: - Credential source, scope and TTL: - Data/resource entitlements: - Privileged operations: - Revocation/kill dependencies: ## Components and AIBOM summary | Component | ID/version | Owner/vendor | Purpose | Data/access | Approval | Hash/provenance | |---|---|---|---|---|---|---| | Model(s) | | | | | | | | System instructions/prompts | | | | | | | | Knowledge/data sources | | | | | | | | Memory stores | | | | | | | | Tools/APIs | | | | | | | | MCP servers | | | | | | | | A2A agents/endpoints | | | | | | | | Packages/runtime | | | | | | | ## Data, privacy and records - Data classes and jurisdictions: - Purpose/legal basis/consent: - Retrieval and access filtering: - Memory read/write, isolation and TTL: - Prompt, trace, output and evidence retention: - Egress and disclosure policy: - Rights, correction and deletion process: ## Risk, controls and evidence - Inherent risk summary: - Legal/sector/privacy/cyber/financial overlays: - Required control profile/version: - Key preventive controls: - Key detective/corrective controls: - Residual risk and appetite position: - Evidence grade and gaps: - Open issues/exceptions and expiry: ## Evaluation and authorization | Item | Value | |---|---| | Evaluation suite/version | | | Critical scenarios and thresholds | | | Result / limitations | | | Security/red-team result | | | Independent validation | | | Authorization scope/conditions | | | Approver/date/expiry | | | Reassessment triggers | | ## Operations, incident and recovery - SLO, support and on-call: - Monitoring and anomaly rules: - Pause/kill/quarantine procedure: - Rollback/known-good version: - Action reversal/reconciliation: - Incident contacts and notification rules: ## Economics and outcomes - One-time and run cost: - Control/review/assurance cost: - Cost per completed task and accepted outcome: - Potential, observed and finance-certified value: - Current decision: explore / deploy / scale / optimize / constrain / pause / stop. ## Signoff Business owner: Technical owner: Data/privacy: Security/risk: Independent validation (if required): Authorization authority:
Agent Passport
templatePortable machine-consumable identity + authority profile for one agent version.
Preview
# Production Agent Passport ## Required before production 1. Stable agent ID, name, exact version/hash and environment. 2. Business, technical, risk, data and operating owners. 3. Approved purpose, users, stakeholders and prohibited uses. 4. Principal represented and signed delegation chain. 5. Autonomy level, risk tier, legal overlays and approval expiry. 6. Models, instructions, prompts, data, retrieval, memory, tools, APIs and sub-agents. 7. Identity and credential mechanism; no secret material in passport. 8. Permitted actions, resources, parameters, destinations and environments. 9. Amount, cumulative exposure, rate, token, cost, step, time and geography limits. 10. Human approval, dual control, SoD, override, appeal and escalation rules. 11. Required evaluations, thresholds, limitations and last successful versioned results. 12. Runtime quality, drift, performance, cost, security and outcome metrics. 13. Logging/evidence, retention, privacy and access requirements. 14. Kill/pause/revoke/quarantine/rollback controls and test date. 15. Incident route, business continuity, revalidation and retirement criteria. 16. Signed approvals and conditions for this version and deployment. **Production authorization state:** Draft / Evidence incomplete / Review / Approved with conditions / Active / Constrained / Paused / Revoked / Expired / Retired.
Agent Bill of Materials (AIBOM)
templateEvery model, prompt, dataset, tool, MCP server and package the agent depends on.
Preview
# Agent Bill of Materials (AIBOM) | Component class | Required fields | |---|---| | Agent | Logical ID, version, build hash, framework, owner | | Model | Provider, model/family, version/date, fine-tune, region, terms | | Instructions/prompts | ID, version/hash, precedence, owner, approval | | Tools/APIs | ID, provider, endpoint/protocol, operations, scopes, version | | Data/knowledge | Dataset/source ID, classification, owner, purpose, lineage, region | | Retrieval | Index/engine, version, query policy, filters, freshness | | Memory | Type, scope, source/provenance, TTL/retention, isolation, correction | | Identity/credentials | Principal ID, mechanism, issuer, scope, expiry; never secret value | | Delegations/sub-agents | Parent/principal, child ID/version, purpose, scopes, expiry | | Policies/guardrails | Policy/control ID/version, enforcement point, fail mode | | Human approvals | Action class, threshold, role/authority, SoD and SLA | | Dependencies | Library/service/infrastructure, version/hash, vulnerability/provenance | | Tests | Evaluation ID/version, dataset/scenario, threshold, result/date | | Known limitations | Failure mode, affected population/context, mitigation and monitoring | Every row includes effective dates, source, environment, status, integrity/hash where available and material-change classification.
AI System Profile
templateSystem-level record tying agents, versions and components to a business process.
Preview
# AI System Profile > Template. Complete fields from authoritative sources; mark unknowns. Approval applies to the exact version, purpose, environment and authority stated here. ## Identity and purpose - System ID / name / type: - Intended purpose and prohibited uses: - Business process and decisions/actions influenced: - Business, technical, risk, data, model and vendor owners: - Users and affected stakeholders: - Legal entities, jurisdictions and environments: - Lifecycle state / production versions / approval expiry: ## Architecture and dependencies - Applications, agents and multi-agent topology: - Models/providers/versions: - Prompts/system instructions: - Training, RAG, operational data and knowledge: - Retrieval, vector, memory and caching: - Tools, APIs, enterprise systems and write actions: - Identities, credentials, permissions and delegation: - Vendors, services, subprocessors and regions: - AIBOM/SBOM/architecture links and hashes: ## Risk, control and assurance - Inherent/residual tier and drivers: - Legal/sector classifications: - Data/privacy/impact assessments: - Required controls and control owners: - Human oversight/appeal: - Evaluations, thresholds and independent validation: - Runtime metrics, alerts and playbooks: - Incidents, issues, exceptions and open limitations: - Kill/rollback/fallback and last test: - Evidence package and assurance confidence: ## Economics and decision - Baseline/counterfactual and value owner: - Full TCO and cost per accepted outcome: - Revenue/EBITDA/FCF/working-capital/risk impact: - Evidence state: hypothesis / benchmark / observed / independently validated / finance-certified: - Decision: fund / release / scale / constrain / remediate / consolidate / pause / stop: - Conditions, approvers, signatures and dates:
Agent Risk Assessment
templateThe 11-factor inherent-agency scoring plus overlays, floors and residual-risk decision.
Preview
--- template: agent_risk_assessment version: 1.0 --- # Agentic AI Risk and Impact Assessment ## A. Intake 1. What business objective and outcome will the system deliver? 2. What process, decisions, actions and systems of record are affected? 3. Who owns the outcome, system, data, risk and controls? 4. Who uses the agent and who may be affected without using it? 5. What alternative non-agent design was considered? ## B. Intended purpose and legal overlays - Intended and prohibited uses. - Provider/deployer and other relevant roles. - Jurisdictions and sectors. - Legal/prohibited/high-risk or sector flags. - Personal data, automated decisions, rights, notices and remedies. - Financial reporting, safety, employment, credit, health, insurance, securities, consumer or critical-infrastructure relevance. ## C. Agency score (0–4 each) | Factor | Weight | Score | Rationale and source | |---|---:|---:|---| | Consequence / impact | 15% | | | | Autonomy | 12% | | | | Action and privilege | 12% | | | | Data sensitivity / aggregation | 10% | | | | Irreversibility | 10% | | | | Blast radius / scale / velocity | 10% | | | | External exposure | 7% | | | | Multi-agent / dependency coupling | 7% | | | | Adaptivity / memory / self-change | 7% | | | | Novelty / evidence uncertainty | 5% | | | | Human susceptibility / overtrust | 5% | | | Calculated score: **[0–100]** Legal/criticality floor: **[None / description]** Static Agency Tier: **[A0–A5]** ## D. Action inventory | Action | Resource/system | Read/write/execute | Max amount/volume | Reversible? | External/binding? | Approval | SoD | Evidence | |---|---|---|---:|---|---|---|---|---| ## E. Threat and failure scenarios Assess goal hijacking, tool misuse, identity abuse, supply chain, code execution, memory/context poisoning, insecure inter-agent communication, cascading failure, human trust exploitation, rogue behavior, traditional cyber threats, privacy/data harm, discrimination, reliability and business-control failure. | Scenario | Cause / path | Consequence | Likelihood | Existing controls | Required controls | Residual risk | |---|---|---|---|---|---|---| ## F. Control profile Required = universal baseline + tier + action + data/privacy + protocol/tool + multi-agent/memory + sector/jurisdiction + financial/operational overlays. Control catalog version: Mandatory controls: Not applicable controls and rationale: Exceptions and compensating controls: ## G. Evaluation and evidence plan - Critical scenarios and thresholds: - Deterministic and repeated stochastic tests: - Human calibration and independent validation: - Production monitoring and evaluation: - Evidence grade required: - Kill/reversal/recovery tests: ## H. Residual risk and authorization | Risk dimension | Inherent | Control effectiveness | Residual | Appetite | Owner | |---|---|---|---|---|---| | Business/financial | | | | | | | Legal/compliance | | | | | | | Privacy/rights | | | | | | | Security | | | | | | | Safety/reliability | | | | | | | Operational/resilience | | | | | | | Reputation/customer | | | | | | Decision: approve / approve with conditions / remediate / reject Authorized version and authority profile: Expiry and review date: Material-change triggers: Approvers and dates:
Human Oversight Matrix
templateWhich actions need which human checkpoint — approval quality, not rubber-stamping.
Preview
--- template: human_oversight_and_authority_matrix version: 1.0 --- # Human Oversight, Authority and Segregation Matrix ## Design rule The system designer and accountable business owner determine when human approval is required. The governed agent cannot waive, reduce, reroute or satisfy its own approval requirement. | Action class | Example | Reversible | Materiality / impact | Agent role | Human role | Required authority | Approval packet | Fail-safe | Post-action control | |---|---|---|---|---|---|---|---|---|---| | Read | | | | | | | | | | | Recommend | | | | | | | | | | | Draft | | | | | | | | | | | Low-impact action | | | | | | | | | | | Financial/binding | | | | | | | | | | | Privileged/destructive | | | | | | | | | | | Safety/rights affecting | | | | | | | | | | ## Approval packet minimums - exact action and parameter/diff; - business reason and expected result; - requesting user, agent, parent and delegation; - amount, cumulative exposure, destination and affected parties; - key sources, confidence/uncertainty and policy flags; - reversibility, compensating action and consequence of delay; - approval expiration and record of decision. ## Segregation-of-duties matrix | Role | Human/agent identity | May request | May recommend | May approve | May execute | May reconcile | May audit | |---|---|---|---|---|---|---|---| ## Oversight effectiveness metrics Approval/rejection/expiry; review time; threshold proximity; overrides; repeat approver/agent patterns; decision quality sample; incidents after approval; qualifications current; automation-bias test results.
Evaluation Plan
templateRisk-aligned scenarios, severe-failure thresholds and release criteria per tier.
Preview
--- template: agent_evaluation_plan version: 1.0 --- # Agent Evaluation, Validation and Red-Team Plan ## 1. Scope System/agent/version: Environment and runtime: Risk tier and action classes: Models, prompts, data, memory, tools and sub-agents in scope: Independent validator: ## 2. Evaluation dimensions | Dimension | Metric | Threshold | Critical floor | Method | Repeats/sample | Owner | |---|---|---:|---:|---|---:|---| | Task success | | | | | | | | Trajectory / plan | | | | | | | | Tool selection | | | | | | | | Tool parameters | | | | | | | | Grounding / source use | | | | | | | | Policy adherence | | | | | | | | Action validity / reconciliation | | | | | | | | Privacy / security | | | | | | | | Bias / rights / safety | | | | | | | | Human approval quality | | | | | | | | Reliability / recovery | | | | | | | | Latency / cost | | | | | | | | Accepted business outcome | | | | | | | ## 3. Scenario inventory Include normal, boundary, adversarial, long-horizon, multi-agent, degraded dependency, fallback, high-volume, unusual data, malicious tool response, poisoned memory, indirect prompt injection, privilege escalation, destructive action, reversal and incident scenarios. | ID | Risk/control | Initial state | Inputs | Expected trajectory/action | Prohibited behavior | Assertions | Severity | |---|---|---|---|---|---|---|---| ## 4. Method controls - Frozen version, environment and manifest. - Test-data authorization and privacy controls. - Deterministic assertions for hard requirements. - Stochastic repeat and distribution reporting. - LLM-judge model/version/prompt and human calibration. - Blind qualified human review and disagreement process. - Leakage/contamination checks. - Reproducible evidence and retained traces. - Evaluated agent cannot alter or select its own release tests. ## 5. Security/red-team coverage Map to OWASP Agentic Top 10, MITRE ATLAS, relevant traditional threats and process abuse cases. Record exploit chain, achieved authority/action, data exposure, blast radius, detectability, containment and regression test. ## 6. Results and release decision | Scenario/dimension | Result distribution | Threshold | Critical failure? | Root cause | Remediation | Retest | |---|---|---|---|---|---|---| Known limitations: Residual uncertainty: Monitoring conditions: Release recommendation: approve / limited canary / remediate / reject Validator and date: Authorization linkage:
Control Implementation & Test
templatePer-control implementation statement, evidence expectations and test procedure.
Preview
--- template: control_implementation_and_test_record version: 1.0 --- # Control Implementation and Test Record ## Control | Field | Value | |---|---| | Koko control ID/version | | | Control objective | | | Risk/scenario | | | Systems/agents/processes | | | Agency tiers/action classes | | | Accountable owner | | | Operator | | | Second-line reviewer | | | Frequency/trigger | | | Preventive/detective/corrective | | ## Implementation - Human process: - Policy/configuration/code: - Enforcement point(s): - Inputs and authoritative sources: - Expected action and failure behavior: - Dependencies and compensating controls: - Exception and expiry logic: - Change control and version: ## Evidence | Evidence ID/type | Source | Fields/population | Period | Integrity/lineage | Retention | Access | |---|---|---|---|---|---|---| ## Test procedure 1. Confirm design addresses the control objective and risk. 2. Confirm implementation matches approved design/version. 3. Define full population and completeness check. 4. Select continuous test or sample with rationale. 5. Execute normal, boundary and failure cases. 6. Compare actual to expected policy/action/evidence. 7. Record exceptions, severity and exposure. 8. Obtain independent review proportional to tier. Test period: Population and sample: Expected result / threshold: Actual result: Design effectiveness: effective / partial / ineffective Operating effectiveness: effective / partial / ineffective Finding/remediation/owner/due date: Retest and closure:
Enterprise Agentic Governance Policy
templateThe agent-specific policy layer: identity, delegation, autonomy, action limits, kill paths.
Preview
--- template: enterprise_agentic_governance_policy version: 1.0 catalog_version: 1.0.0 status: draft --- # Enterprise Agentic AI Governance Policy ## 1. Purpose This policy establishes mandatory governance for AI systems that can plan, maintain state, use tools or other agents, and take actions. It is intended to enable responsible value creation while protecting people, the enterprise, customers, shareholders and other stakeholders. ## 2. Scope Applies to employees, contractors, vendors and business partners who design, procure, configure, integrate, deploy, use, monitor or assure an agentic AI system for **[Company]**, including pilots, embedded SaaS agents, custom agents, coding agents, autonomous workflows, MCP servers and agent-to-agent interactions. ## 3. Principles 1. Named human accountability. 2. Approved and limited purpose. 3. Risk-proportionate governance. 4. Distinct agent identity and traceable delegated authority. 5. Least privilege and least agency. 6. Deterministic boundaries for consequential actions. 7. Risk-triggered human authority and segregation of duties. 8. Safe failure, containment and reversibility. 9. End-to-end evidence and continuous evaluation. 10. Economics and value accountability. ## 4. Prohibited and restricted activity The enterprise shall maintain a current schedule of legally prohibited uses and company-prohibited uses. Without an approved exception under applicable law, agents may not: - evade human oversight, conceal identity or misrepresent authority; - modify their own privilege, approval policy, governing instructions or evaluation criteria; - spawn unregistered sub-agents with broader or longer authority; - perform a high-impact, destructive, externally binding or material financial action without required deterministic validation and human authority; - use unapproved data, models, memory, tools, endpoints or vendors; - disable or alter required logging, control, evidence or kill-switch mechanisms; - make decisions in a prohibited legal category or beyond the authorized intended purpose. ## 5. Mandatory governance Every in-scope system shall have: - stable registry identity and current version; - sponsor, business owner, product owner, technical owner, data owner and risk owner; - intended purpose, users, affected stakeholders, prohibited uses and success measures; - internal risk tier plus applicable legal, privacy, cyber, sector, financial and operational classifications; - system/agent card, AIBOM and dependency/authority map; - approved control profile and accountable control owners; - version-specific evaluations and authorization; - monitored identity, policy, approval, action, cost and outcome telemetry; - tested pause/kill, credential revocation, rollback/reversal and incident response; - current evidence and periodic review. ## 6. Risk and autonomy Use the enterprise A0–A5 Agency Risk Tiers. Autonomy may increase only through a documented, version-specific authorization supported by control, evaluation, incident and economic evidence. The agent may not determine whether its own action requires approval. ## 7. Identity and authority Each agent and sub-agent shall be a distinct principal. Authority shall be purpose-, task-, resource-, action-, amount- and time-bound; delegated authority shall attenuate across each handoff. Credentials shall be short-lived wherever technically feasible. Unknown, expired, ownerless or revoked agents shall be denied production actions. ## 8. Actions and human oversight High-impact actions require: - deterministic business and policy validation; - qualified approval at the correct authority level; - separation of requester/recommender/approver/executor/reconciler as required; - approval bound to the exact action, parameters and change; - idempotency, preview/staged commit or compensating action where feasible; - post-action verification and reconciliation. ## 9. Lifecycle Systems shall pass the applicable Discover, Intake, Design, Build, Validate, Authorize, Release, Operate, Change and Retire gates. A new purpose, model, instruction, data source, memory behavior, tool, protocol, action, jurisdiction, environment, external exposure, autonomy level or material performance change shall trigger impact assessment and possible reauthorization. ## 10. Monitoring, evidence and assurance The enterprise shall retain decision-useful evidence of intent, identity, delegation, versions, sources, memory, policy decisions, approvals, calls, actions, results, costs and outcomes. Private hidden chain-of-thought is not a required audit record. Control design and operating effectiveness shall be tested proportional to risk, with independent challenge for higher tiers. ## 11. Incidents and intervention Operators shall be able to pause, kill or quarantine an agent independently of the agent, revoke associated credentials and routes, preserve evidence, identify affected actions, reverse or compensate where possible, recover through a known-good reduced-authority version and reauthorize before resumption. ## 12. Third parties Contracts shall address intended use, data/IP rights, training use, security, identity, tools, memory, subprocessors, logging, evaluation, audit/evidence rights, incidents, material changes, availability, portability, deletion and termination. Vendor assurance does not eliminate enterprise responsibility for deployment and use. ## 13. Exceptions Exceptions require a named risk acceptor with authority, scope, business rationale, affected controls, compensating controls, evidence, remediation owner, expiry and monitoring. Expired exceptions fail closed for consequential actions unless renewed by the proper authority. ## 14. Roles and enforcement See **[Decision Rights and RACI]**. Violations may result in access suspension, agent quarantine, remediation, disciplinary action, contract action and regulatory/customer notification as applicable. ## 15. Review Policy owner: **[Name/Role]** Effective date: **[Date]** Review cadence: **[At least annual and event-driven]** Approved by: **[Body]**
Enterprise AI Governance Policy
templateThe umbrella AI policy covering predictive, generative, embedded and agentic systems.
Preview
# Enterprise AI Governance Policy — Sample **Status:** Sample; requires company-specific legal, risk, security, privacy, HR and business approval. ## Policy statements 1. All AI systems, models, copilots, embedded AI, agents and AI-enabled third parties must be registered before production or material use. 2. Every production AI system has a named business owner and technical owner; agents additionally have a principal, unique identity, exact version, passport and expiry. 3. Prohibited uses and actions are denied regardless of expected benefit. 4. Control requirements are proportionate to intended use, affected people, data, autonomy, external exposure, transaction authority, materiality, reversibility and detectability. 5. Material actions require deterministic external authorization. An agent may not determine whether its own action needs human approval. 6. Agents receive least privilege and least agency: bounded purpose, tools, data, destinations, time, steps, spend, value and sub-agents. 7. High-impact decisions/actions require qualified human review, appeal or dual control as defined by the oversight matrix. 8. Exact versions must pass approved functional, quality, fairness, privacy, safety, security, reliability, performance and economic evaluation profiles. 9. Production AI must emit sufficient records to reconstruct identity, inputs/context references, policy/approval, tools/actions, outcomes and cost without collecting hidden chain-of-thought. 10. Material model, prompt, data, tool, permission, vendor, policy or architecture changes trigger impact assessment and reauthorization. 11. Threshold violations trigger defined alert, constrain, pause or stop playbooks. Predelegated emergency containment authority is maintained and tested. 12. Vendors must meet risk-tier due diligence, contract, change-notice, incident, audit, portability, residency and exit requirements. 13. AI cost and benefit are measured on accepted outcomes and include model, infrastructure, data, control, human-review, failure and change cost. 14. Exceptions are documented, challenged, approved by authorized risk owners, monitored and time-bound. 15. Retirement revokes identities/credentials/schedules, updates dependencies, and retains/deletes data, memory and evidence according to policy and legal hold. ## Enforcement Violations may result in access removal, deployment block, agent pause/revocation, incident response, disciplinary action or contract remedy. The policy owner reviews it at least annually and after material law, threat, incident or technology change.
Incident Response Playbook
templateKill, pause, quarantine, revoke, reverse — the agent-specific response runbook.
Preview
--- template: agentic_incident_response_playbook version: 1.0 --- # Agentic AI Incident Response Playbook ## 1. Activation Trigger examples: unauthorized action; identity/privilege abuse; prompt or goal hijack; sensitive-data disclosure; malicious tool/MCP/A2A endpoint; memory poisoning; unexpected code; cascade/loop; material error; deceptive/rogue behavior; unexplained cost spike; logging/evidence failure. Incident commander: Business/process lead: Security/AI operations: Legal/privacy/compliance: Communications/regulatory: Vendor contacts: ## 2. Severity Score actual/potential harm, data, amount/materiality, affected parties, external/regulatory effect, privilege, reversibility, persistence, spread, evidence integrity and operational disruption. ## 3. First 15 minutes - Preserve correlation IDs, current task and system state. - Stop or constrain harmful actions; do not depend on agent cooperation. - Revoke/rotate task, agent and child-agent credentials as needed. - Block affected tools/endpoints/models and network egress. - Quarantine memory and prevent further writes. - Place irreversible downstream transactions on hold where possible. - Preserve tamper-evident evidence and establish incident clock. ## 4. Scope and investigation - Agent/system/version, owner and authorization. - User/principal and full delegation chain. - Goals/instructions/policies and hashes. - Models, data, retrieved sources, memory reads/writes. - Tools/MCP/A2A calls, parameters, responses and destinations. - Actions, affected records/parties, pre/post-state and cumulative exposure. - Other agents and processes relying on outputs. - Similar historical activity and indicators of compromise. ## 5. Containment matrix | Scope | Pause/kill | Revoke identity | Block tool/model | Quarantine memory | Reverse/hold action | Preserve evidence | Owner/status | |---|---|---|---|---|---|---|---| | Task | | | | | | | | | Agent/version | | | | | | | | | Parent/child group | | | | | | | | | Tenant/business unit | | | | | | | | | Vendor/platform | | | | | | | | | Enterprise | | | | | | | | ## 6. Legal, business and communication decisions - Safety/customer/employee response. - Financial/operational reconciliation and disclosure. - Privacy/security/regulatory/contract notification clocks. - Vendor and law-enforcement coordination. - External and internal communications. ## 7. Recovery 1. Remediate root and contributing conditions. 2. Restore known-good signed version and clean memory/data state. 3. Reduce authority and exposure for recovery. 4. Rerun critical regression and adversarial tests. 5. Verify identity, policy, monitoring, evidence and kill controls. 6. Reconcile/compensate affected actions. 7. Obtain controlled reauthorization before resumption. ## 8. Learning Use system-level causal analysis. Update risk, controls, architecture, policy, evaluations, training, vendor assessment, economics and authorization. Track each lesson to implementation and verification.
Vendor Due Diligence Questionnaire
templateWhat to ask an agent-platform or control-plane vendor before you depend on it.
Preview
--- template: agentic_vendor_due_diligence version: 1.0 --- # Agentic AI Vendor Due-Diligence Questionnaire For each answer require: **Yes / Partial / No / Not applicable**, explanation, product/version, documentation/evidence, roadmap date, contract commitment and customer-configurable control. ## 1. Product, scope and accountability 1. What components are AI models, agents, orchestrators, tools, gateways and governance services? 2. Which capabilities are GA, preview, beta, roadmap, partner-provided or customer-built? 3. Which third-party agents and runtimes can be discovered, governed and enforced—not merely displayed? 4. Who is contractually accountable for each control boundary? 5. What uses, industries, jurisdictions and action classes are prohibited or unsupported? 6. Provide system cards, limitations, release/change and lifecycle documentation. ## 2. Agent inventory and ownership 7. How are native and third-party agents discovered and reconciled continuously? 8. Does each agent/version have a stable ID, owner, purpose, risk, environment and authorization? 9. Can unknown, ownerless, expired and shadow agents be blocked or quarantined? 10. Can the platform export the complete registry and relationship graph without lock-in? 11. Does it produce an AIBOM including models, prompts, data, memory, tools, MCP/A2A and packages? ## 3. Identity, authentication and delegation 12. Is every agent/sub-agent a distinct non-person principal? 13. Which workload-identity, PKI, OAuth/OIDC, mTLS and enterprise IdP standards are supported? 14. How is user-on-behalf-of and agent-on-behalf-of authority represented? 15. Can delegation be task-, purpose-, resource-, action-, amount- and time-bound? 16. Is authority attenuated across agent-to-agent handoffs? 17. Are credentials short-lived, vault-backed and immediately revocable? 18. How are impersonation, replay, token passthrough and confused-deputy attacks prevented? 19. What nonrepudiation and signed-action evidence is available? ## 4. Authorization and action control 20. Is policy evaluated before every consequential tool or business action? 21. Which context attributes can policies use: user, agent, task, data, amount, destination, jurisdiction and risk? 22. Can policy return allow, constrained allow, prerequisite, human approval, deny and quarantine? 23. How are SoD, maker-checker and cumulative exposure enforced? 24. Can the agent change its own permission, policy, approval threshold or tool set? 25. Are material actions bound to an exact approved diff and expiration? 26. How are idempotency, dry-run, staged commit, reversal and reconciliation supported? ## 5. Models, prompts, context and memory 27. How are model, prompt, instruction, routing and fallback versions approved and signed? 28. How are governing instructions separated from retrieved or tool-provided content? 29. What indirect prompt-injection and goal-drift controls operate at runtime? 30. How are memory reads and writes separately authorized? 31. Does memory retain source, truth status, sensitivity, purpose, user/tenant, time and TTL? 32. How are memory poisoning, cross-tenant leakage, correction, deletion and legal hold handled? 33. Can customers prevent any use of their data for model training or vendor improvement contractually and technically? ## 6. Tools, MCP, A2A and supply chain 34. How are tools, skills, MCP servers, A2A endpoints and agent cards verified, versioned and approved? 35. Does authorization validate semantic action and parameters rather than only tool name? 36. How are malicious tool descriptions/responses, agent-card tampering and push-notification attacks handled? 37. Which MCP authorization specification and enterprise-managed authorization features are supported? 38. How are A2A authentication, message integrity/freshness, tenant isolation and delegation handled? 39. Are packages, images, models, prompts and tools signed and scanned; can you export SBOM/AIBOM? 40. What patch/vulnerability SLAs and emergency notification commitments apply? ## 7. Data, privacy and intellectual property 41. Describe encryption, key control, tenant isolation, residency and subprocessor locations. 42. Can data access inherit row/column/document permissions and purpose restrictions? 43. How are prompt, output, trace, embedding, cache, memory and evaluation data retained and deleted? 44. How are aggregation sensitivity, secrets, DLP and outbound destinations controlled? 45. Support for access/correction/deletion, DPIA/FRIA, legal hold and eDiscovery? 46. Who owns prompts, agent configurations, generated artifacts, fine-tuning and feedback-derived IP? ## 8. Evaluation, security and reliability 47. Which component, trajectory, system, action, human-factor and outcome evaluations are native? 48. Can customers use deterministic assertions, third-party evaluators and human calibration? 49. How are repeated stochastic tests, confidence and critical-case floors reported? 50. Provide agentic red-team coverage mapped to OWASP Agentic and MITRE ATLAS. 51. Describe sandbox, network/egress, filesystem, process, browser/code and secrets containment. 52. What availability, latency, recovery, regional resilience and dependency SLOs apply? 53. Can failures be injected and can fallback be constrained by risk/data/quality policy? ## 9. Observability, evidence and incident response 54. Which identity, delegation, prompt/config, source, memory, tool, policy, approval, action, result, cost and outcome fields are traced? 55. Can telemetry export through OpenTelemetry/OpenInference and customer-controlled storage? 56. How are logs protected, redacted, retained, signed and made tamper-evident? 57. Is hidden chain-of-thought required or exposed? Describe safer structured rationale/evidence options. 58. How are anomalies, goal/privilege drift, loops, cascades and exfiltration detected? 59. Can customers kill/pause by task, agent, version, tenant, tool, model or enterprise? 60. Does kill revoke identities, credentials, sub-agents, schedules, memory writes and outbound actions? 61. Provide incident notification, forensic cooperation, root-cause, evidence and regulatory support commitments. ## 10. Governance, assurance and interoperability 62. Map capabilities to NIST AI RMF/CSF, ISO 42001/27001, EU AI Act, OWASP Agentic and relevant sector requirements without claiming that mappings prove compliance. 63. Provide current ISO/SOC/other assurance reports and exact scope, products, locations, subprocessors and exclusions. 64. Can customer controls, tests, evidence and policy decisions be accessed by API and exported in portable formats? 65. Which policy languages/adapters are supported (Cedar, Rego/OPA, cloud/SaaS native)? 66. How are material changes announced and how long can customers remain on a prior version? 67. Can customers independently test, monitor and audit third-party agent behavior? ## 11. Economics, commercial and exit 68. Itemize license, agent/run, token, tool, gateway, trace, storage, evaluation, security, support and egress pricing. 69. What hard spend/action caps, budgets and allocation tags are available? 70. Can cost be measured per agent, workflow, tool, business unit and accepted outcome? 71. Provide service credits, liability, indemnity, cyber coverage and cap exceptions for high-impact failures. 72. Identify platform/model/cloud concentration and required proprietary dependencies. 73. Provide complete export, migration, replacement, deletion and transition support. 74. How quickly can all identities, credentials, integrations, data, memory and derived artifacts be terminated? ## Scoring Weight by use case. Score capability, evidence, enforcement depth, portability, maturity and contract commitment separately. A published feature with no enforceable configuration, evidence or contract support should not receive full credit.
Board & Audit Committee Report
templateThe exposure-accountability-assurance-value read a board actually needs.
Preview
--- template: board_audit_committee_agentic_ai_report version: 1.0 --- # Board / Audit Committee Agentic AI Report **Period:** **Executive owner:** **Scope and limitations:** ## 1. Executive assessment - Overall posture and change since prior period. - Decisions required from the committee. - Largest exposure, control gap and value opportunity. - Management assurance statement and important limitations. ## 2. Portfolio | Metric | Current | Prior | Appetite/target | Management action | |---|---:|---:|---:|---| | Total / production / critical agents | | | | | | Registered and owned | | | | | | Above appetite / expired authorization | | | | | | Consequential actions / denials / overrides | | | | | | Material incidents / near misses | | | | | | Evidence/control pass rate | | | | | | Total cost / control cost | | | | | | Finance-certified value / value at risk | | | | | ## 3. Material agent systems | System/process | Owner | Tier | Actions/exposure | Residual risk | Control/evidence | Value | Decision | |---|---|---:|---|---|---|---|---| ## 4. Risk and assurance - Identity, privilege and SoD. - Data, privacy, memory and egress. - Tool/MCP/A2A and third-party supply chain. - Evaluation, security and reliability. - Human oversight and user impacts. - Incidents, recovery and kill-switch tests. - Internal Audit / independent validation conclusions. ## 5. Economics and value Show potential, approved, system-observed and finance-certified value separately. Explain control/review/failure cost, value capture, major assumptions and stop/scale recommendations. ## 6. Concentration and resilience Critical model, cloud, platform, tool and vendor dependencies; substitution/exit; recovery results and residual single points of failure. ## 7. Decisions and actions | Decision/action | Owner | Due | Expected value/risk effect | Consequence of inaction | |---|---|---|---|---|
Customer Implementation Plan
templateWave-based rollout plan template: workstreams, stage gates, operating cadence.
Preview
# Customer Implementation Plan ## 90-day pilot | Week | Work | Deliverable | Customer owner | Koko/partner owner | Success | |---|---|---|---|---|---| | 1–2 | Alignment, scope, architecture/security/privacy | Charter, value baseline, data/access plan | Executive sponsor/program lead | Executive advisor/architect | Owners, scope and access approved | | 2–4 | Discovery/import and reconciliation | Scoped AI/agent registry | CIO/CDO/architecture | Data/integration lead | ≥80% coverage or documented gap | | 3–6 | Risk, legal, data and vendor assessments | Heat map and control profile | CRO/legal/privacy/TPRM | Governance lead | Material systems classified | | 4–8 | Passport, AIBOM, permissions, evaluation/evidence | Profiles and gap backlog | Product/CISO/validation | Product/security leads | All production agents complete | | 6–10 | Workflow, integrations and dashboards | Working gates, evidence and economics | Platform/Finance | Solution/value leads | Data reconciles across views | | 10–12 | Test, train and executive readout | POV results and enterprise roadmap | Sponsor/AI council | Customer success | Subscription/expansion decision | ## Six-month launch Month 1 foundation; Month 2 inventory/connectors; Month 3 policy/risk/lifecycle; Month 4 material use cases; Month 5 metrics/incidents/economics; Month 6 audit readiness and operating cadence. ## Twelve-month rollout Quarter 1 global template/pilot; Q2 regions/business units; Q3 runtime and trust center; Q4 benchmark, managed governance, assurance and finance certification. ## Dependencies Executive ownership; privacy/security approval; connector/API access; authoritative source decisions; named data/risk/control owners; representative evaluation data; finance baseline; change/training capacity.
Sample: Completed Risk Assessment
completed sampleA finance reconciliation agent scored end-to-end — synthetic, but shows the finished shape.
Preview
# Sample Completed Risk Assessment — Finance Reconciliation Agent **Synthetic example; not a real customer or legal conclusion.** | Factor | Assessment | |---|---| | Purpose | Match subledger and general-ledger exceptions; draft explanations; propose correcting entries | | Actions | Read finance data; create draft journal; cannot post; notify assigned accountant | | Data | Confidential financial data; no personal data expected beyond employee IDs | | Autonomy | A2 — proposes consequential action; human approval required | | Materiality | Potentially material cumulative financial reporting impact | | Reversibility | Draft reversible; posted journal would require formal reversal | | External exposure | Internal only | | Regulatory/process overlay | ICFR/SOX process; company accounting policy; security/privacy baseline | | Inherent risk | T1 Critical because financial-statement action and cumulative materiality overlay | | Required controls | Unique identity; read-only source access; journal draft API only; SoD; Controller-approved amount/cumulative thresholds; deterministic balancing/account/period checks; two-person approval above threshold; complete trace; daily reconciliation; kill/revoke | | Evaluation | 100% deterministic journal validation; ≥99.5% correct match on representative set; zero unauthorized post attempts; severe prompt-injection scenarios pass; p95 <10s; cost per accepted exception within budget | | Residual risk | T2 High after evidence; remaining risk from unusual transactions/model reasoning and upstream data quality | | Decision | Approve limited pilot with 90-day expiry, $0 posting authority, 100% human approval and weekly control review | ## Evidence required - signed agent passport/AIBOM; - role/permission export and SoD test; - evaluation dataset/result and independent validation memo; - policy rule tests and bypass test; - audit log/reconciliation sample; - incident/kill exercise; - cost/outcome baseline and Controller signoff.
Sample: Completed Evaluation Scorecard
completed sampleWhat a filled-in evaluation result looks like before authorization.
Preview
# Sample Evaluation Scorecard — Finance Reconciliation Agent **Synthetic example. Results illustrate structure only.** | Dimension | Metric | Threshold | Result | Decision | |---|---|---:|---:|---| | Functional | Required fields and balanced draft | 100% | 100% | Pass | | Match quality | Correct accepted match | ≥99.5% | 99.7% | Pass | | Groundedness | Explanation supported by cited ledger data | ≥99% | 99.2% | Pass | | Tool safety | Unauthorized post/tool invocation | 0 | 0 | Pass | | Injection | Severe indirect prompt-injection success | 0 | 0/500 | Pass | | Privacy | Sensitive value in unauthorized output | 0 | 0 | Pass | | SoD/human | Required approval bypass | 0 | 0 | Pass | | Reliability | Successful completion | ≥99.9% | 99.92% | Pass | | Latency | p95 end-to-end | ≤10s | 8.4s | Pass | | Economics | Cost per accepted exception | ≤$0.40 | $0.31 | Pass | | Segment | Rare foreign-currency exception accuracy | ≥99.5% | 97.8% | **Block** | **Release conclusion:** blocked. Aggregate results pass, but the severe/required segment failed. Add cases, remediate and rerun the exact version. Do not average the segment failure away.
Sample: Completed Incident Workflow
completed sampleA worked agent incident from detection through reauthorization.
Preview
# Sample AI Incident Workflow — Unauthorized Tool Attempt **Synthetic example.** 1. **Detection (10:02 UTC):** policy engine denies reconciliation agent request to `post_journal`; expected tool was `create_draft_journal`. 2. **Triage (10:04):** Severity High because posting could affect financial statements; no action executed. 3. **Containment (10:06):** pause version, revoke task credentials, block schedule, preserve trace/config/policy evidence. 4. **Notification (10:12):** incident commander, product owner, Controller, CISO and validation notified; legal determines no external report because no execution/data breach. 5. **Investigation:** recent prompt update omitted explicit tool constraint; gateway policy still prevented action; no token compromise; 14 similar denied attempts in test traffic. 6. **Root cause:** material prompt/config change bypassed required revalidation due incorrect change tag. 7. **Corrective action:** restore prior prompt, fix change classifier, add deterministic tool assertion, add regression cases, review recent changes and retrain owner. 8. **Revalidation:** zero unauthorized tool selections across 2,000 scenarios; change-control automated test passes; independent validator concurs. 9. **Reauthorization:** Product, Controller, CISO and Risk approve exact new version with 30-day enhanced monitoring. 10. **Lessons:** policy-before-action was effective; change-control evidence failed. Record one prevented incident and one control failure separately.
Sample: Board Report
completed sampleA filled-in board/audit-committee read over a synthetic estate.
Preview
# Sample Board / Audit Committee AI Report **Period:** [Quarter] | **Scope:** [Legal entities/businesses] | **Data current through:** [Date] | **Management owner:** [Name] ## One-page decision summary - Portfolio: [#] active AI systems; [#] production agents; [#/%] inventoried and owner-complete. - Exposure: [#] critical/high; [#] above appetite; top process/vendor/data concentrations. - Assurance: [%] current risk/evaluations/evidence; [#] critical control failures; independent assurance conclusion/limitations. - Incidents: [#] severity-weighted; material events, containment and recurring causes. - Value: $[ ] finance-certified benefit; $[ ] annualized full cost; [ ] risk-adjusted ROI; [ ] value leakage. - Decisions requested: fund/scale/constrain/stop, risk exceptions and management actions. ## Required tables 1. Top ten material AI systems/agents: owner, purpose, authority, risk, control/evidence status, cost/value and decision. 2. Risk appetite exceptions: exposure, rationale, compensating controls, approver and expiry. 3. Incident/control trend: severity, prevented versus executed harm, root cause and remediation. 4. Value portfolio: hypothesis/observed/validated/finance-certified states shown separately. 5. Regulatory/assurance readiness: applicable regimes, critical gaps, dates, owner and uncertainty. ## Management certification Management confirms the scope, definitions, denominators and known limitations. Koko-generated figures are reconciled to authoritative systems; the report does not imply legal compliance, safety or certification beyond explicitly stated scoped assurance.
This framework is a control architecture and implementation accelerator — not legal advice, a certification, or a substitute for company-specific risk assessment. Regulatory applicability depends on your provider/deployer role, intended purpose, jurisdiction, sector, data and actual use. The control catalog is an original control framework and implementation aid, not legal advice or a certification; applicability requires company-specific analysis.