A Touch Stone Publishers White Paper COO Series | Agentic AI Governance


A high-volume decision stream forks at the point of human sign-off, yet both the reviewed and the unreviewed path arrive at the same visible outcome, until an aggregate harm reveals which path was never watched.

The Finding

An illustrative exposure profile shows why a high-volume workflow should be ranked by decision volume and per-decision exposure before governance work begins.

In many enterprises, the Chief Operating Officer oversees the largest concentration of high-volume agentic workflows. Order fulfillment, customer remediation, scheduling, and quality exception handling are attractive candidates for automation because they contain thousands of small, repeatable decisions. That volume is the COO’s operating advantage. It also increases the importance of the officer-level oversight duty described in In re McDonald’s Corporation Stockholder Derivative Litigation (Del. Ch., January 25, 2023). The court held that corporate officers owe an oversight duty within their areas of responsibility, while emphasizing that liability requires bad faith rather than an imperfect system or an ordinary mistake.

The exposure does not sit where most COOs expect it. It does not sit primarily in whether any single agent makes a bad decision. It sits in the gap between having twenty agentic systems and having documented, for even one of them, who is personally accountable, what that system may decide without a human, and how its behavior actually reaches that accountable person before a material harm accumulates. Of the three architectural gaps every agentic deployment carries, the COO is structurally the officer most likely to be carrying the third and most dangerous one: no working escalation mechanism. Not because operations leaders are careless. Because it is materially harder to document decision rights and escalation paths for twenty operational agents than for one financial system, and most COOs have never been asked to do it.

This is the gap that survives longest undetected. A missing named officer is visible the first time anyone asks who owns a system. A missing decision-rights document is visible the first time anyone asks what the agent is authorized to do. A missing escalation mechanism is invisible by construction, because it only reveals itself during an actual failure, and by then the failure has usually already aggregated past the point where a single decision would have looked like a problem worth raising. This white paper exists to close that specific gap, in the COO’s own operating language, before an aggregate harm closes it for them.


Why Volume Changes the Legal Calculus, Not Just the Operational One

Most executives still treat AI governance as a technology risk: the model will err, the model will drift, the model will occasionally produce a decision a human would not have made. That instinct is a category error for every officer, and it is the most expensive category error a COO specifically can make, because operational scale is exactly the condition under which the error compounds.

Caremark liability does not arise simply because a system makes a mistake. The legal inquiry is far more demanding and turns on bad faith. McDonald’s nevertheless makes the operating question concrete: did the responsible officer make a good-faith effort to establish and monitor an information system reasonably designed for the risks within that officer’s remit? A documented, monitored, appropriately bounded architecture does not create a safe harbor, and an undocumented architecture does not automatically establish liability. The difference is evidentiary. One can show how oversight was designed and exercised. The other leaves the organization trying to reconstruct intent and practice after the event.

Red-flag monitoring becomes less informative when an agent can execute thousands of individually unremarkable decisions that aggregate into material harm. No single refund, scheduling adjustment, or waived exception may appear significant. A systematic pattern across ten thousand decisions can be. This paper therefore focuses on a management requirement, not a new legal test: operations should monitor category-level drift and route defined exceptions to a named decision-maker before an external inquiry reveals the pattern.

The operating choice is not speed or governance. It is whether the organization can scale both at once. A ranked inventory, reusable accountability contract, and tested escalation pattern allow the COO to govern the highest-exposure systems first without freezing lower-risk deployments. That is the practical advantage of an architecture over a one-off remediation.


The Evidence Record the COO’s Function Is Least Prepared to Produce

A formal inquiry tests prompt logs, model version history, and architecture diagrams that many COO inventories have not yet produced.

When litigation, regulatory review, internal audit, or a board inquiry tests an agentic system, the organization may need to reconstruct what the system did, which configuration governed the decision, what monitoring signal appeared, and who acted. A board minute stating that AI governance was discussed cannot answer those questions by itself. A record showing a named officer, a defined decision-rights boundary, a monitoring cadence, and documented intervention can.

For the COO, that evidence requirement can apply across scheduling, fulfillment, quality exception handling, and customer remediation at the same time. Building a bespoke record for every system does not scale. The architecture described here solves that problem through a ranked inventory, a named accountable director for each priority system, and escalation rules designed to detect pattern drift rather than only large individual decisions.


The Accountability Contract Model: The Specific Fix, Named

The Accountability Contract Model's four elements, authority, decision boundary, success metric and escalation threshold, and timeline, only hold if each is filled in for a named, specific agentic system.

The fix for the COO’s exposure is not a policy document, a working group, or a broader AI use standard. It is a named accountability conversation, held with a specific person, about a specific agentic system’s authority, and the Touch Stone Publishers Accountability Contract Model names precisely what that conversation must contain.

The Accountability Contract Model holds that accountability is not a value an organization declares. It is a conversation a leader has, stating clearly what needs to be done, what authority is granted, what success looks like, and what the timeline is, before asking for results. Most organizations skip this conversation entirely and then are surprised when an outcome cannot be defended, because the outcome was never precisely assigned to begin with. Applied to an agentic system, the model requires four elements stated with the same specificity Glenn requires of any accountability conversation, not assumed as understood by the parties involved.

Authority. The named director must have documented authority to configure, suspend, or modify the agentic system’s operating parameters within the boundary the board has set at the policy layer. A director who can be asked about a system but cannot act on it does not hold accountability. They hold exposure without the authority that would let them discharge it.

Decision boundary. The categories of decision the agent may make without further human sign-off must be written down and version-controlled, specific enough that an independent reviewer asking what this agent was authorized to decide gets an artifact, not an institutional memory. For a COO managing volume, this boundary typically needs to be stated in terms of a dollar threshold, a complaint category, or an exception type, not in the general language of “the agent handles routine cases.”

Success metric and escalation threshold. What “working correctly” means for this system must be defined in advance, in terms specific enough to test, alongside the defined deviation that triggers escalation to the named director. This is the element many operations functions have never stated in writing, because informal deployment treats “the agent seems to be working” as sufficient. An independent review quickly exposes that gap.

Timeline. The cadence on which the contract is reviewed, and the date of the next scheduled review, must be stated, not left to whenever someone next thinks to check.

An organization that has assigned an agentic system to “the operations team” or “the automation working group” has not built an accountability contract. It has assigned collective activity without individual decision authority. The Accountability Contract Model converts undocumented operational risks into named, bounded, reviewable accountability relationships.


Worked Scenario: The Refund Concierge Agent, Before and After

The same category-level decision stream took five months to reveal a drift under an aggregate dashboard alone, and three to four weeks once a category-drift escalation rule was overlaid.

The following scenario is illustrative. It combines common operating conditions to show how a category-level control behaves; it does not describe a named enforcement matter.

Consider a customer remediation function that deployed an agent, internally called the Refund Concierge, to auto-approve refunds and account credits below a $150 threshold. The deployment appeared successful. Average resolution time for eligible complaints fell from four days to under two hours, and the team treated the rollout as complete once the approval rate stabilized.

Before. No document existed defining which categories of complaint the agent was authorized to evaluate, only an informal understanding that it handled “the straightforward ones.” No named director held documented authority to modify the agent’s approval criteria; the function that had built the deployment, a three-person automation team reporting through a shared operations manager, treated ongoing tuning of the model’s pattern-matching as a routine engineering task rather than a governance decision. Over roughly five months, the agent’s pattern-matching began systematically under-crediting complaints in a specific category, billing disputes tied to a subscription renewal flow the company had changed earlier that year, at a rate the operations team’s own dashboards would have shown as anomalous if anyone had been comparing categories against each other rather than watching only the aggregate approval rate, which stayed within its normal range throughout. Each individual under-credit was too small, typically eleven to twenty dollars below what a human reviewer would have approved, to trigger a manual review on its own, and no single decision looked, in isolation, like anything worth escalating. The pattern was invisible to the people running the agent and invisible to the COO, because nothing in the architecture was built to compare category-level behavior against a baseline. It surfaced only when a state consumer protection office’s routine sampling of resolved complaints flagged the category, at which point the aggregate harm had already accumulated across several thousand affected accounts, and the operations team’s first task was reconstructing, after the fact, a decision history the logging retention policy had not been built to preserve.

After. The COO required a documented escalation rule specific to this failure mode: any category of decision where the agent’s approval rate for a defined complaint type deviated more than a stated percentage, in this instance twelve percentage points, from that category’s trailing-quarter baseline, would surface automatically to a named Director of Customer Operations for review within one business day, independent of whether the agent’s aggregate approval rate looked normal. The director held documented authority to suspend the agent’s authority over the flagged category, require manual processing of new complaints in that category, and adjust the underlying pattern-matching criteria, all without waiting for a broader review cycle. The same underlying model drift, tested retroactively against the historical data from the prior deployment, would have surfaced within three to four weeks of onset rather than five months, entirely inside the organization’s own architecture, reviewed and resolved by a named director whose intervention was documented in writing, rather than discovered by an external regulator’s sample and reconstructed under the pressure of an active inquiry. The dollar difference between those two timelines, in remediation cost, regulatory exposure, and the operations team’s own credibility with the COO’s office, is the entire argument for building the escalation mechanism before it is needed rather than after.

The lesson this scenario carries for every other agentic system in the COO’s inventory is specific: an escalation mechanism built around single-decision thresholds alone will miss exactly the failure mode most likely to occur at operational scale, because operational agents are designed to make individually small decisions in high volume. The mechanism that catches aggregate harm has to compare a category’s current behavior against its own recent baseline, continuously, not merely flag decisions that cross an absolute dollar or severity line.


Building an Escalation Mechanism for Volume, Not for Individual Decisions

The distinction the Refund Concierge scenario makes concrete is the one most operations functions get backwards when they first attempt this work. A single-decision threshold, an alert when any one refund exceeds a stated dollar amount, is easy to build and catches almost none of the failure modes that actually threaten a high-volume agentic deployment. The failure modes that matter at the COO’s scale are pattern failures: a category of decision, a class of customer, a type of exception, drifting away from its own historical baseline while every individual decision inside that drift stays small enough to look unremarkable.

The escalation architecture that closes this gap has three components, and all three need to exist before the mechanism can be said to function rather than merely to exist on paper. First, a defined baseline for each meaningful category of decision the agent handles, built from a trailing period long enough to smooth normal variation, typically a full quarter. Second, a stated deviation threshold, expressed as a percentage move away from that baseline, that triggers mandatory escalation regardless of whether the aggregate approval rate across all categories looks normal. Third, a named director, specific to the system, with documented authority to act on the escalation the moment it fires, not merely to be informed of it. A mechanism missing any one of these three components will look, from the outside, indistinguishable from one that actually works, right up until the moment it is tested by a real failure.

This is also where the escalation mechanism connects upward. The named director’s intervention record, including a simulated test when no real event has occurred, is what the COO reports to the CEO and what reaches the board’s review of its most exposed agentic systems. A configured alert is not enough. The organization should demonstrate that the alert reaches the right person and that the person can exercise the stated intervention rights.


Ranking the Inventory: Where the COO Starts

The scale problem this white paper has described, twenty agentic systems where a CFO might have one, means the COO cannot document everything at once and should not try to. The organizing discipline is a ranked inventory: every agentic system in operations, ordered by decision volume and dollar exposure per decision, so that the systems most likely to produce an aggregate harm before anyone notices are the systems that get a documented accountability contract and a tested escalation mechanism first.

This ranking is not the same as ranking by the size of any single decision the system makes. A system that occasionally approves a large individual transaction but does so rarely, with each instance already subject to manual review by convention, may carry less aggregate exposure than a system making thousands of small decisions a day with no category-level comparison built in. The Refund Concierge scenario is the second pattern, not the first, and it is the pattern most operations leaders underweight, because the instinct to worry about the large individual decision is natural and the instinct to worry about the small decision made ten thousand times is not.

For each system that reaches the top of this ranked list, the work is the same four-part build described above: name the director, document the decision boundary in writing, build and test the category-drift escalation rule, and record the first review with the CEO. Systems further down the list can carry lighter documentation initially, a boundary statement and a named owner without a fully tested drift mechanism, provided the ranking itself is revisited on a defined schedule so that a system’s exposure is not assumed to be permanently low simply because it was not near the top of the list at the last review.


What This Looks Like Thirty Days In

A COO who has not yet built this architecture does not need a multi-quarter program to establish the foundation. In the first week, the highest-value action is the ranked inventory itself: every agentic system in operations, ordered by decision volume and dollar exposure per decision, with the top five identified honestly, including the systems the COO’s own team may prefer not to examine closely because they have been running quietly and successfully for years. In the second week, the top-ranked system receives a named accountable director and a written decision-rights boundary specific enough that an outside reviewer could test whether a given decision fell inside or outside it. In the third week, the operations team builds the category-drift escalation rule for that system, appropriate to catching pattern drift rather than only single large decisions, and tests it against historical data or a simulated deviation to confirm it actually triggers. In the fourth week, the COO reviews the first cycle of escalation data with the named director, including any interventions the rule produced, real or simulated, and reports the result to the CEO. That report, however brief, is the first demonstrated instance of the architecture functioning, and it is the specific artifact that separates a documented but untested policy from a working oversight system.

None of this requires slowing operational deployment. It requires running the accountability build alongside deployment, system by system, starting with the systems where volume makes the invisible aggregate harm most likely, rather than waiting for a governance program to mature before the first system is properly documented.


What This White Paper Connects To

The ranked inventory, named director, Accountability Contract Model, and category-drift escalation mechanism are developed in the Executive Leadership Playbook on Agentic AI Governance. The Playbook sets out the board policy layer, officer accountability layer, and technical execution layer. It also includes the ten-dimension Governance Readiness Diagnostic used to identify the next action for each system.


The Legacy Test

A ranked inventory ordered by decision volume and dollar exposure shows exactly which systems still lack a named director, a documented boundary, or a tested escalation mechanism.

The measure of the COO who builds this architecture is not whether every agent performed well during the officer’s tenure. It is whether the escalation mechanism, named directors, and documented decision boundaries still work for the successor who inherits the seat. The COO who builds this before an incident acts from conviction. The one who builds it after an inquiry begins is responding to evidence the organization should already have been able to produce.


Sources and Scope

  1. Delaware Court of Chancery, In re McDonald’s Corporation Stockholder Derivative Litigation, C.A. No. 2021-0324-JTL (January 25, 2023): https://law.justia.com/cases/delaware/court-of-chancery/2023/c-a-no-2021-0324-jtl.html
  2. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework 1.0 and Playbook. The framework is voluntary and calls for documented roles, inventories, monitoring, and human oversight: https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
  3. U.S. Department of Justice, Evaluation of Corporate Compliance Programs (September 2024), including governance of emerging technologies and testing whether controls work in practice: https://www.justice.gov/criminal/criminal-fraud/page/file/937501/dl

Legal authorities establish duties and enforcement facts. The ranked inventory, Accountability Contract Model, category-drift thresholds, review cadence, and 30-day sequence are Touch Stone Publishers recommendations. This paper is governance analysis, not legal advice.

Choose your next step

Continue the collection or save the topic. Return to the Agentic AI Governance collection and choose the paper most relevant to your current decision.

Continue or save this topic

Stay current. Receive new research and executive intelligence as it is published.

Join the newsletter

Go deeper. The annual intelligence membership provides the governed research, implementation detail, and continuing updates behind the public library.

Request a membership seat