Skip to main content
    All shows

    Monday, August 31 · 30 min

    Stakeholder Communication, Lifecycle & Operational Enablement

    0:00-:--
    Speed

    Transcript

    Koko: Every answer in these two domains is an artifact created before the moment it is needed, in a form somebody else can check. That is the one instinct. If you leave this module with nothing else, leave with that.

    Sam: And I want to explain why that instinct is hard-won. I have watched a sound AI program lose its executive sponsor — not because the technology failed, but because the steering group could not point to a document that said who had approved what, when. Everything that should have existed beforehand was being reconstructed after the fact. That is a very different conversation to be in.

    Koko: That is exactly the failure mode these two domains are designed to prevent. So let's lay out what we're working with. Domain six is stakeholder communication and lifecycle management — fourteen percent of the Professional exam. Domain seven is developer productivity and operational enablement — seven percent. Together that's twenty-one percent, the largest single slice in this series. And here is the thing the prep materials mostly skip: there is no Foundations equivalent for this material. Foundations asks whether you can build. This asks whether the organization can trust you afterward.

    Sam: The framing I use is that domain six is what an outsider is told, and domain seven is what an operator can see. Two audiences, one underlying question.

    Koko: That is a clean way to hold it. And once you see it that way, the same instinct covers all fourteen ideas in both domains. Artifact, before the moment, checkable by someone else. Now — the first artifact in the sequence is structured discovery. Here is the fact, stated plainly: structured discovery is a written instrument, not a conversation you hope goes well. It is a fixed question set where each question is explicitly attached to the decision it unblocks, and you run it identically with everyone.

    Sam: The identically part is doing real work there. If the question changes depending on who is in the room, the answers are not comparable, and you end up with a patchwork of informal agreements instead of a single auditable record.

    Koko: Exactly. Now the scenario. You have twelve questions, grouped by the decision each one settles. One group settles where data may be processed. Another settles what may be stored and for how long. Another establishes what runs unattended and what requires human sign-off. One question names the approver — not the role, the person by name. And then the highest-yield question in the whole instrument: what was tried before this, and why did it stop.

    Sam: That last one surfaces the hidden constraints. If a previous vendor relationship ended for compliance reasons, or an internal prototype was shelved after a data incident, that history is live constraint material. You want it in week one, not week eight.

    Koko: Right. So here is the analogy. Structured discovery is a surveyor's report before you buy the house. A surveyor does not ask what you would like the kitchen to look like. A surveyor tells you which wall is load-bearing. One of those conversations changes what you can build; the other one just feels productive.

    Sam: And the load-bearing wall you miss in discovery is the one you find when demolition is already underway.

    Koko: Which brings us to the takeaway. Ask what the system may not do before you ask what it should do. Prohibitions delete entire architectures. Wishes almost never do. If data residency rules out a cloud region, three months of design work goes with it. If a retention policy prohibits storing conversation logs, your observability stack needs to be rebuilt from scratch. Those constraints need to arrive first.

    Sam: The discriminator I press on in reviews is whether each question on the instrument is tied to a specific downstream decision. If I can't name the decision a question unblocks, the question probably belongs in a product roadmap meeting, not in discovery.

    Koko: That is the test. And now the trap, and I want to name it clearly because it is very easy to mistake for rigor. The misconception is the feature questionnaire — the one that asks stakeholders to rate capabilities one to five, maybe ten or fifteen items, covers a lot of ground, feels thorough. The problem is it contains no constraints. It is entirely about what people want. A data residency rule does not show up on a feature questionnaire. It shows up at the security review, which is long after your architecture options stopped being cheap.

    Sam: The version of this I keep seeing on teams is treating the requirements workshop and the constraint discovery as the same meeting. They are not. One asks what the system should do; the other asks what it cannot do. Running them together means the wants crowd out the prohibitions every time.

    Koko: Two different meetings, two different instruments, two different outputs. That is the first artifact. Everything else in these domains builds on having this one done before anything else moves.

    Koko: So there are two halves to one job here. The first half: communicating a decision is not announcing the winner. It is recording what was chosen, against what alternatives, under which constraints, at what cost, and crucially — what would make you revisit it. The second half: a trade-off explained in your units is not communicated at all. A CFO owns cost and risk. A head of operations owns throughput and exceptions. A compliance officer owns evidence and auditability. If you hand them your units, you have done the analysis but not the communication.

    Sam: The framing I use on my teams is: the record and the table are two separate artifacts with two separate jobs. The record is the institutional memory. The table is the translation layer.

    Koko: Exactly right, and the exam will test whether you know what goes in each. The record is one page, five headings: context; options considered, each with what it buys and costs; the decision in one sentence with a date and a named decider; consequences; and a revisit trigger — something like a volume threshold that, when crossed, puts the decision back on the table.

    Sam: The revisit trigger is the piece most teams skip. They treat the decision as permanent. What is the structure on the table side?

    Koko: Three columns, one row per option. The option; what it buys, expressed in the reader's own metric — dollars per month, cases per hour, whatever they track; and what it costs in that same metric. Same row, same units, both directions.

    Sam: So for a CFO the cost column might be dollars per month in additional human review. For an operations lead it might be average handling time per exception.

    Koko: Right. Now the analogy for the record: it is the written judgment behind a verdict. The verdict says who won. The judgment says why — so the next case can argue against the reasoning instead of relitigating the facts. Without the judgment, every future conversation restarts from zero.

    Sam: And six months later someone walks in proposing the rejected option, with no idea it was already on the table.

    Koko: Which is the first trap, and we will get there. The analogy for the table is a doctor explaining a prognosis. Same clinical fact, three conversations — to the patient, to the family, to the insurer — each in different terms. A surgeon who lists only the benefits has obtained agreement, not consent. That distinction matters on the exam.

    Sam: Agreement is downstream of the upside. Consent requires the downside to land in terms the person actually owns.

    Koko: Good restatement. The takeaway, and there are two. Write the record while the rejected options are still fresh, and name the decider — a decision with no owner cannot be reopened by anyone, because there is no one to go back to. Then, on the table: write the cost column first. If you cannot fill it, the analysis is unfinished. Not unpresentable — unfinished.

    Sam: The cost column as a completeness check. I like that. The option sitting next to this on the exam will be 'present the winning option with its benefits highlighted,' which approves faster and feels cleaner.

    Koko: That is the benefits slide — three options, upside only, one highlighted. It does approve faster. And it turns every later cost into a surprise. That is how a sound design loses its sponsor: not because the design was wrong, but because the sponsor never owned the cost side and has no language for it when it arrives.

    Sam: The version of this I keep seeing on teams is recording only the chosen option and its rationale. The rejected alternatives just disappear. Then six months on someone proposes one of them again, nobody remembers why it was rejected, and you relitigate the whole thing.

    Koko: That is the first trap exactly. Recording only the chosen option and its benefits. The fix is mechanical: the record has one row per option considered, not one row for the winner. Name what each one cost, name who decided, name what would change the answer. That is the whole job.

    Koko: Two documents, one underlying question in different units: what has this organization agreed to accept? That question gets asked twice before launch — once about availability and latency, once about accuracy. And neither answer is real unless it is written down with enough specificity to be enforceable.

    Sam: The way I frame it on teams: a service-level statement is not a service-level agreement until every row has four things — the metric, the target, how and over what window it is measured, and what happens when you breach it. Miss any one of those and you have a wish, not a contract.

    Koko: Exactly right. And the accuracy statement is the same structure applied to model quality. Ninety-two percent accuracy is excellent or unacceptable depending entirely on what was agreed about the other eight percent. The number means nothing without its context.

    Sam: The worked example I keep coming back to: one row in the SLA — accuracy, ninety-two percent on the frozen regression set, measured per release, breach blocks the release. That is four fields, all four populated. And beside it sits a one-page accuracy statement, signed before build.

    Koko: And the accuracy statement does something the SLA row cannot do on its own — it separates the two error types. Because a false positive and a false negative do not cost the same thing in most businesses. You have to name them separately and get agreement on what each one costs.

    Sam: Right. And then you translate the residual into concrete volume — not 'eight percent error rate' but forty cases a month landing on a named team. That is the number someone has to operationally absorb, and it needs to be accepted in writing. The accuracy statement also carries the pause threshold — the point at which observed error in production triggers a stop.

    Koko: Here is the analogy that makes this stick. There is a difference between someone saying 'we will get it to you soon' and handing you a docket with a date, a measurement, and a consequence if it slips. Those are not the same kind of object. One is a reassurance. The other is a receipt.

    Sam: The receipt is what you can point to six months later when the conversation starts drifting.

    Koko: The instinct to carry in: derive every target from something actually measured, and get the residual accepted in writing while everyone is still optimistic. Optimism is a resource. Use it to get signatures, not to skip the document.

    Sam: The trap the exam puts next to this looks like a real SLA. It will read something like 'ninety-nine point nine percent uptime, sub-second response' — and stop there. No measurement window, no stated consequence, no acknowledgment of the dependencies the team actually controls. It has the right vocabulary and none of the structure.

    Koko: And the second trap is the live demo treated as the commitment. The room watches six curated inputs run cleanly, calibrates on that, and from that point on the first real error in production reads as a defect — because nobody signed off on what the actual error rate would be at scale. The demo set the expectation, and that expectation was never formally corrected.

    Sam: Which is why the accuracy statement has to exist as a separate signed artifact before build. The demo cannot substitute for it, no matter how well the demo went.

    Koko: Here is the fact, and the second half is what makes the first half load-bearing. A handoff is complete when somebody who was not in the design can operate the system at three in the morning. An architecture diagram is not that document.

    Sam: And that second part is the one teams miss. The diagram tells you what was built. It does not tell you which dependency goes down first under load, or who to call when it does.

    Koko: Exactly. And the same logic applies to phases. Teams report a phase as a status — we are in build, we are in handoff — when a phase is actually defined by the artifact that lets you leave it. You do not exit handoff because time passed. You exit it because a specific document exists and the receiving team has rehearsed against it.

    Sam: So the gate table makes that concrete. Discovery closes on the constraint set and a measured baseline. Build closes on the evaluation set passing its thresholds. Handoff closes on the runbook, rehearsed by the receiving team. Monitoring does not close at all — that row names a review cadence instead.

    Koko: Right. And the runbook itself has seven sections. What the system does and does not do. Its dependencies and their failure behavior. Failure modes observed so far, with symptoms and fixes. How to pause it and who is authorized to. Where the logs and cost figures live. The evaluation set and its thresholds. And the named owner for configuration and tool permissions.

    Sam: That last one matters more than it looks. If nobody's name is on configuration permissions, the first incident becomes a permissions hunt at midnight.

    Koko: The analogy I want you to carry is handing over the keys, not the drawings. Think of a boiler manual: which fuse trips when the kettle and the heater run together, and the number to ring at three in the morning. That is the document. The architectural schematics of the boiler are not.

    Sam: And for the phase gate, it is a driving test, not a driving course. You leave the learner stage because an examiner signed something. Not because you had enough lessons.

    Koko: That is a cleaner version of the same instinct. So here is the takeaway to carry into the exam. Write the runbook from the on-call person's questions, not from the design. Record what normal looks like. Report against artifacts, not percentages. And know that the most dangerous moment in the lifecycle is the transition into operation — because the people who will run the system never chose what running it takes.

    Sam: That last point is where I see the most incidents in practice. The team that built it understood every edge case. The team that inherited it is meeting those edge cases for the first time, alone, at night.

    Koko: Now the traps, because the exam puts two right next to this. First: handing over the diagram and the repository. Neither document says what to do when the system behaves oddly. So the first incident escalates straight back to the people who built it. That is not a handoff. That is a delayed dependency.

    Sam: The version of this I keep seeing on teams is that the runbook gets written after the first incident — because that is when the gaps become visible. Which means the handoff was never actually closed.

    Koko: And the second trap is percent complete. The exam will offer it as a progress metric and it will look reasonable. It is unfalsifiable, and it hides which phase was skipped. That is how a program sits at ninety percent for two months. The artifact either exists and has been signed or it has not. That is the only falsifiable status.

    Koko: Deprecation is the one lifecycle event that has published mechanics rather than craft. Models move through four states: active, legacy, deprecated, and retired. When a model hits deprecated, it keeps working — requests still go through — but there is a retirement date attached, and after that date, requests fail. Anthropic commits to notifying customers with active deployments at least sixty days before retiring a publicly released model. Sixty days, on the calendar, in writing.

    Sam: That sixty-day floor is the thing people miss. It's not sixty days from when they happen to check the changelog — it's sixty days from when the notice goes out to accounts with active deployments.

    Koko: Exactly. So the window is real, but only if you're watching for the notice. Now here's the scenario: picture a migration notice sent the week the deprecation is announced — not the week before the lights go out. That notice names the retirement date, lists every affected surface with its owner, shows the replacement model's results against the existing frozen regression set, and documents the cutover plan and the rollback path. That is what good looks like.

    Sam: The regression set piece is the one I push on teams hardest. You can't just say the replacement is better in general — you have to show it on your eval set, the one frozen before the migration, so you have an apples-to-apples comparison your stakeholders can actually read.

    Koko: Right, and that's why the takeaway has three parts: make deprecation notices a named responsibility — one person owns the inbox, not the whole team — export usage by API key and by model so you know exactly which surfaces are affected, and re-run the frozen evaluation set on the replacement before you move anything. The analogy I like for this is a lease with a notice period. The letter arrives months ahead. The building works perfectly until the date. The tenant who filed the letter unread is the one moving out over a weekend.

    Sam: That's a good one because the building doesn't degrade — there's no warning wobble before the cutoff. It just stops accepting your key on the retirement date.

    Koko: None. So the trap here is treating deprecation as an engineering ticket — doing the swap quietly, validating it internally, and only telling business owners after the model has already changed underneath them. The misconception is that 'we validated it' is sufficient communication. A business owner who was never told the model changed, and then sees output drift in production, does not experience that as a successful migration. The stakeholder event is not the cutover. The stakeholder event starts the day the notice lands.

    Sam: The version of this I keep seeing on teams is the model swap goes into a deploy commit with no corresponding stakeholder update, because everyone assumed someone else sent the communication. Named responsibility closes that gap.

    Koko: That's the whole thing. Sixty days is enough runway. The question the exam is testing is whether you treat that runway as an engineering timeline or a communication timeline — and the answer is both, with the communication starting first.

    Koko: Every configuration decision is secretly an access-control decision. Settings in these systems live at different scopes — organization, workspace, project, individual — and the scope is what determines who can override a value and who can audit it. Productivity comes second. The scope question comes first.

    Sam: That framing is exactly right and it changes how you architect a rollout. The thing I see platform teams get backwards is they think about what the setting does before they think about who owns it.

    Koko: Right. And Skills fit the same logic. A Skill packages a procedure — instructions and executable code — as a directory the model loads on demand. So the question isn't just 'what does this Skill do,' it's 'at what scope does it live and who controls that scope.' One more thing: custom Skills do not sync between surfaces. If you build one for the desktop client, it is not automatically available in the web surface.

    Sam: The sync behavior trips people up in practice. We had a Skill working perfectly in one context and the team assumed it was everywhere. It was not.

    Koko: Classic. Now let me make this concrete. A platform team commits the shared code-review procedure as a project-scoped Skill inside the repository. Every engineer on that project gets exactly one version. Changing it requires a pull request and goes through code review — the same gate as any other code. Meanwhile key issuance and spend limits sit at the workspace scope, so only the workspace admin can touch those. Two different scopes, two different populations of people who can make a change.

    Sam: What's the discriminator for which scope you choose? Because project and workspace can both feel like 'shared.'

    Koko: The discriminator is: does the setting follow the person or the project? Spend limits follow the workspace regardless of which project is open. The review procedure follows the repository regardless of which engineer is in the seat. If you can answer that question, the scope choice follows.

    Sam: And the description field on a Skill — I tell my teams that is not a comment for humans browsing a list.

    Koko: Exactly. A Skill's description is trigger metadata. It's written in the phrases your team actually uses, because that is how the model decides whether to load the Skill for a given task. The description is the only part of a Skill that is always in context. The rest loads on demand.

    Koko: Here is the analogy I use. Think of the recipe pinned on the kitchen wall versus the chef's personal notebook. The pinned copy is what makes two different shifts produce the same dish. A Skill from an outside source is an ingredient from a supplier — you audit suppliers before you put their ingredients on the line.

    Sam: That audit instinct applies directly to third-party Skills. You don't install them the way you'd install a browser extension without reading the permissions.

    Koko: Good. So the takeaway to carry into the exam room: for every setting, ask who can change it, and ask whether it follows the person or the project. Those two questions surface the right scope every time.

    Koko: Now the traps, and there are two. First misconception: that standardizing means writing a document telling everyone to configure things the same way. A document is a suggestion. Configuration not enforced at a scope that somebody administers is a suggestion. A project-scoped Skill is enforced. A wiki page is not.

    Sam: The exam version of this will offer 'create documentation for the team' as a standardization answer and it will look totally reasonable.

    Koko: It will look very reasonable. Second misconception: that a Skill is documentation. A Skill gives the model instructions and executable code. It is not a README. If the exam offers an answer that treats Skills as a place to store reference material for humans, that is wrong. The model runs the Skill. Humans read the description.

    Koko: Last one. And this one ties the whole domain together. Telemetry. Here is the fact: observability is three independent signals with separate switches. Metrics give you token and cost counters. Log events capture prompts and tool results. And traces — currently in beta — give you a span per interaction, per model request, per tool call, with subagent spans nesting under the parent's. Three switches. You can have one on and the other two dark.

    Sam: And the identity piece is what catches people. Every event names the credential by default — so you see your service account on every row, not the end user, unless you explicitly inject the user identity at request time.

    Koko: Exactly. And cost works the same way — reportable by key, by workspace, by model, but only across dimensions that existed at request time. You cannot slice data by a dimension you never attached.

    Sam: The scenario from practice that makes this concrete: a team reports the agent is slow. They pull the trace. The model request spans are fast. What is fat is a single tool span, and inside it there is a permission-wait child span that is eating most of the latency. The bottleneck is a human approval queue. Not the model at all.

    Koko: And you only see that if the trace was on. If you only had metrics, you would see elevated end-to-end latency and you would be looking at the wrong thing. That same team, by the way, gave each product its own workspace and each environment its own key — so their cost data is recoverable by product and by environment.

    Sam: The analogy I keep using for this is a flight data recorder. It captures aircraft parameters by default. The cockpit voice channel only records if that channel was fitted. It names the aircraft, not the pilot. And it cannot be installed after the accident.

    Koko: That is the whole thing in one image. The takeaway: turn telemetry on at the first deployment and verify data is arriving, because export failures are silent. Decide attribution and the cost boundary before launch. After launch is too late to recover what you did not collect.

    Sam: Now the traps, because there are four of them clustered here and the exam will use all of them. The first one I see on teams constantly: the assumption that telemetry captures what the agent did. It is structural by default. Content — the actual prompts, the tool results — is opt-in.

    Koko: Second trap: treating per-user attribution as a dashboard nicety. It is not cosmetic. A service-account trail can only answer the question the application did something. An audit trail needs to answer who. That is the difference between a log and an audit trail.

    Sam: Third: the belief that granularity can be added later. One shared key across four products makes per-product cost permanently unrecoverable. The option next to the right answer on the exam will say you can backfill with tags — you cannot, if the key boundary was never there.

    Koko: And the big one. A bad answer means a bad prompt. Most failures live outside the model — in tool execution, in permission waits, in downstream API responses. The trace is the only thing that shows you where the time actually went.

    Koko: Alright. Seven modules. Let Sam close it, because he has shipped these things.

    Sam: The shape is the same across all fourteen ideas in this domain. The constraint set, the decision record, the trade-off table, the accuracy statement, the runbook, the gate table, the migration notice, the enforced scope, the telemetry, the open trace — every one of them is created before the moment it is needed.

    Koko: That is the instinct the exam rewards. When two options both look defensible, the right one holds a document somebody outside the team can check.

    Sam: Stakeholder cannot read the code. They can read the trade-off table. The auditor cannot observe the model. They can read the accuracy statement. The on-call engineer at two in the morning cannot reconstruct your intent. They can read the runbook.

    Koko: Every choice in this domain is really the same choice: did you make the decision legible before the moment of accountability arrived? If yes, you pass the exam and you build something your team can actually operate.

    Sam: The drills and flashcards for everything we covered across the whole domain are at KokoAI Academy — koko knows dot A I.

    Koko: Go build something, and go pass this exam. You have done the work.