A regional health system approves an AI-assisted triage tool for its telehealth intake line, with one standing condition: every high-acuity case flagged by the system gets a licensed nurse's review before a callback disposition is finalized. Eighteen months later, the model still performs exactly as validated. Nobody has touched it. The workflow around it, meanwhile, has moved four separate times, through four separate decisions, none of which anyone thought to call a change to an AI-governed system.
Attrition costs the intake line two of five nurses in a single quarter. A supervisor, watching a callback-time metric slide toward a missed service target, authorizes what she logs as a staffing accommodation: only the two highest acuity tiers out of five will get mandatory nurse review; the other three will be dispositioned by non-clinical staff following the tool's suggested category directly. She plans to restore full review once hiring catches up. Hiring does not catch up. Six months later, a cost initiative run by a different manager benchmarks staffing against the now-reduced review load and concludes current staffing is appropriately sized, quietly making the exception permanent without anyone in that review knowing a full-review baseline had ever existed. A vendor contract renewal swaps the escalation contact without updating the internal runbook. Fourteen months in, the workflow the tool actually runs inside bears almost no resemblance to the one the quality committee approved, and there is no record anywhere that would let a director reconstruct that fact in an afternoon.
This is the pattern hiding inside most companies running AI in a regulated process. Model degradation is the risk everyone is watching for: it is rare, monitored, and usually caught by the vendor or the function that validated it. Workflow degradation is common, unmonitored by anyone with board-reporting responsibility, and caused entirely by ordinary operating decisions: a headcount reduction, a reorganization, a productivity target, a shift-coverage fix made at eleven at night on a Friday. None of those decisions touch the model. Any one of them can quietly convert a board-approved, human-supervised system into something the board never approved and does not know exists.
The audit that checked the wrong population came back clean because it was designed to
In the health system's case, two safeguards existed and both missed it, for a specific and instructive reason. The vendor's monthly dashboard reported a stable ninety-four percent "flagged-case disposition compliance" rate throughout the entire period, a metric measuring only whether a flagged case received some disposition inside the contracted window, never whether that disposition involved a nurse. The number looked identical before and after the staffing change because it was never built to distinguish the two states. Separately, an internal quality audit eleven months in sampled call recordings for accuracy, but drew its sample only from the two highest-acuity tiers, the sole tiers still receiving mandatory review. The audit came back clean on the narrow question it happened to ask. The population where the drift had actually occurred was never in the sampling frame, not because anyone hid it, but because the audit assumed the approved workflow was still the operating one.
No Delaware court had decided an AI-specific Caremark claim as of this writing. Applying existing board-oversight doctrine to a case like this is reasoned legal analysis, not settled law, and that limitation should be stated plainly rather than talked around. What existing doctrine establishes with more confidence is the shape of the theory itself, and the theory does not turn on whether the model worked. It turns on whether a reasonable system existed to tell the board that the conditions under which it approved the system had changed. A supervisor who narrows a review requirement to solve a real staffing problem, with no channel that surfaces the change to whoever approved the original workflow, has built the exact fact pattern this doctrine is designed to catch, independent of whether a single patient was ever harmed.
The operating floor has no natural owner, and every other function assumes someone else has it
General counsel owns legal exposure. Compliance owns the regulatory map. The chief information officer owns technical assurance. The Chief Operating Officer owns the one thing none of the other three can see from where they sit: the daily operating reality of the workflow itself, who is actually reviewing, how fast, what gets skipped when the queue backs up, and what happens when the one person who understood the escalation path leaves the company. That operating floor is where AI systems actually fail in production, and it is also, in most companies, the layer with the least formal reporting obligation attached to it.
The honest objection here deserves a real answer rather than a wave of the hand. A COO can reasonably argue that staffing levels, review cadence, and escalation routing are management prerogative, and that inviting board-level scrutiny into every operational adjustment would bury the function in governance overhead it was never built to carry. That objection is correct about a real cost. Boards are not equipped to review staffing decisions in real time, and a governance model that required prior sign-off on every workflow tweak would slow the business for no protective benefit that justifies it.
The objection is right about authority and wrong about reporting
This is where the argument crystallizes, and it has a name: the Governance Boundary Principle. The board governs. Management manages. Neither one is supposed to run the other's job, and the specific failure this case describes is not a board crossing into operations, it is the reverse blind spot nobody names as clearly: an operating function quietly absorbing a reporting obligation that was never its call to make unilaterally. A COO retains full authority to run the operation. What a COO does not retain, once a workflow touches a system the board has already flagged as material and regulated, is the unilateral right to be the only person in the building who knows a review step was removed. The authority is operational. The obligation to report a defined, narrow category of change is fiduciary. Confusing the two, in either direction, is the boundary failure.
A workable version of this asks for almost nothing that would slow an operation down. It names five or six trigger conditions per material system, tied to a number operations already tracks for other reasons, such as review coverage falling below a stated percentage of its approved baseline, rather than to a manager's self-report of intent. A manager can mislabel a decision as a temporary staffing fix. A manager cannot make the coverage percentage itself say something other than what it is, once the number is the trigger and not the label attached to it. That single design choice, trigger on the metric rather than the manager's framing, is what keeps the whole system from becoming an honor code.
What a COO can put in place in thirty days
The first step is an inventory: within two weeks, name every AI-touched workflow under COO authority that also touches a regulated asset, patient care, grid infrastructure, or a permitted process, and attach the original approval conditions to each one, the named review roles, staffing ratios, and escalation contacts. The second step turns that inventory into triggers: for each material workflow, draft the specific operating condition, in language operations already uses, that would mean the workflow has moved outside what was approved, and name one person, not a department, accountable for reporting it. The third step is a template, not a policy: a short log where a manager records what happened, when, who was told, and what the plan is to restore the original baseline, built to be filled out inside minutes rather than reconstructed months later from memory and email. The fourth step is a calendar entry: confirm the board or committee cadence that already exists has a standing line item for this workflow inventory, and get the first placement on the schedule within the same thirty days.
One escalation condition belongs on that calendar permanently. Any material workflow that loses more than a fifth of its approved human-review coverage, for any operational reason, triggers notification to the named owner and a log entry within five business days, regardless of whether the reduction was framed internally as temporary, departmental, or unrelated to the system it touches.
This analysis is developed at greater length in The Fiduciary Vacuum Executive Leadership Playbook, which includes a companion white paper for the Chief Operating Officer on approval drift and the reporting standard this piece describes.
The Close
The Legacy Test for a COO on this question has two clauses, and both have to hold at once. First: could this operation run efficiently, absorb staffing and cost pressure, and make the ordinary calls operations exists to make, without waiting on board sign-off for every adjustment. Second: if a regulator, a plaintiff's expert, or a successor COO reconstructed this workflow's history eighteen months from now, would the record show that the trigger conditions and the reporting habit were named in advance, months before any inquiry arrived, rather than assembled after the fact once a subpoena made the question urgent. A COO who satisfies only the first clause has an efficient operation and an exposed board. A COO who satisfies both has done the job the role actually requires: not simply running the workflow well, but building the record before it was needed, not in response to the moment it became necessary.
