Every Chief Risk Officer who presents a clean quarterly risk report to a board committee this year is making an implicit claim: that the organization's risk exposure has been measured, not merely observed. For agentic AI systems, that claim is usually false, and it is false in a specific, mechanical way that has nothing to do with how carefully any single decision was reviewed.
The Delaware Court of Chancery held in McDonald's (January 26, 2023) that corporate officers, not directors alone, carry a personal fiduciary duty of oversight. That doctrine is not new law in 2026. What is new is a fact pattern the doctrine has never been tested against: autonomous systems making thousands of individually unremarkable decisions, each one inside its approved threshold, each one defensible on its own, while the pattern across those decisions drifts steadily into an aggregate position the organization would never have approved if anyone had measured it directly.
This is not a hypothetical edge case. It is the specific, structural failure mode agentic AI introduces into risk management, and it is different in kind from the failure modes most risk taxonomies were built to catch.
The Mechanism Is Aggregation, Not Error
A risk appetite threshold, in its classical form, is built to catch a transaction or a position that crosses a defined line on the day it happens. That model assumes the unit worth watching is the individual decision, because for most of the history of enterprise risk management, the individual decision was where the exposure lived. A single large loan, a single large trade, a single large operational loss: these were catchable precisely because they were visible, discrete, and large enough on their own to warrant scrutiny.
Agentic systems break that assumption at the root. An agent making thousands of individually small, individually compliant decisions does not generate a single event large enough to trip a transaction-level threshold. It generates a pattern, visible only when the decisions are aggregated and measured against a baseline, not against a static limit. A credit decisioning agent that denies applications from a specific demographic proxy at a rate gradually increasing over a quarter never denies any single application in a way that looks, on its own, like a fair lending violation. The violation exists only in the aggregate, and a threshold built to catch single transactions will never see it, no matter how conscientiously it is monitored.
The traditional Caremark defense relies on the existence of red flags: evidence that a risk was visible and a reasonable officer would have noticed it. Frontier AI systems do not reliably produce red flags in a form a human would recognize before a failure occurs. For a risk officer, the defense that has protected risk functions for decades, that monitoring did not surface a problem because no single event crossed a threshold, is no longer a defense. It is closer to an admission that the monitoring architecture was built for a risk profile agentic systems do not have.
The Governance Boundary Principle, Applied to Risk Appetite
The Governance Boundary Principle holds that the board governs and management manages, and that an organization begins to fail, quietly at first and then suddenly, the moment either crosses into the other's territory. For the Chief Risk Officer, that principle draws the exact line that determines whether a risk architecture is defensible or merely decorative.
The board's role is to set risk appetite: the categories of aggregate exposure the organization will and will not tolerate, stated at the policy level and documented in committee charter language specific enough that a director could locate it without asking management first. The board does not approve individual agent configurations or review individual escalation events. The CRO's role is the layer beneath that boundary: translating the board's risk appetite policy into named, system-specific thresholds and aggregation rules for every agentic system operating inside a mission-critical function.
A board with an excellent AI risk appetite policy sitting on top of a CRO who has not translated that policy into system-specific thresholds has a functioning board layer over a nonfunctioning management layer. McDonald's evaluates the officer's own conduct, not the board's, and a strong board layer does not protect a weak officer layer because the two are assessed separately.
The Register That Passed Every Test and Missed Everything
A risk committee at a mid-market financial services firm received a quarterly report containing exactly one line, reading "AI and automation risk: monitored, no material issues." That single line covered eleven distinct agentic systems spread across three business units: a credit decisioning agent, a claims adjudication agent, a procurement approval agent, a customer remediation agent, and seven others spanning scheduling, quality exception handling, and vendor onboarding.
None of the eleven systems was named individually in the register. None had an individually assigned accountable officer. No aggregation rule existed for any of them, which meant that even the business units informally tracking single-decision limits had no mechanism for detecting whether a pattern of individually compliant decisions was drifting into an aggregate position the organization would not have approved if asked directly. The report's single line was true in the narrowest sense: no decision had crossed an existing alert threshold. It was false in every sense that mattered to a committee doing its job, because it told the committee nothing about which of the eleven systems was closest to an unacceptable aggregate position.
The CRO who corrected this required each of the eleven systems entered individually into the register, each with its own named accountable officer and its own documented threshold, set using the organization's existing risk appetite framework rather than a new framework built from scratch. Each entry also received an aggregation rule measuring a rate over a rolling ninety-day window against the system's own trailing baseline, escalating automatically to the named officer if the rate drifted beyond a defined band, regardless of whether any single decision had crossed the existing dollar-value threshold.
The next quarterly report named each system individually, with its threshold status and its aggregation reading stated specifically. The committee's question changed, in the space of one quarter, from whether AI risk was under control, a question with no specific answer available, to which of the eleven systems was closest to its threshold, a question with a checkable answer the CRO could produce on request.
What the Board Should Ask For, and What It Should Not
A risk committee should hear which named agentic systems are closest to their thresholds, whether any aggregation rule fired in the reporting period, and whether the named accountable officer for any system demonstrably intervened. It should not be asked, and should not volunteer, to approve the specific numeric threshold set for any individual system or review the underlying configuration of any individual agent. That level of specificity belongs to the CRO and the named officers beneath the CRO, operating inside the risk appetite category the board has already set at the policy level.
A committee that starts reviewing individual thresholds has drifted into officer-level risk management while believing it is exercising additional diligence. The opposite is closer to the truth: a board reviewing individual thresholds has less capacity, not more, to ask the one question that actually tests the architecture, which is whether the register as a whole shows a functioning aggregation mechanism rather than a static list of individually approved numbers.
Some risk leaders will argue that a quarterly committee review, conducted by people with deep institutional judgment, already catches what an automated aggregation rule would catch. That argument holds for risks that announce themselves through a single visible event. It does not hold for a pattern spread across thousands of small decisions, none of which is individually reviewable in a quarterly cycle, and none of which a human reviewer would flag without already knowing where to look. The committee's judgment is not the failure point. The absence of a mechanism that tells the committee where to direct that judgment is.
At the next risk committee meeting, ask for a list of every production agentic system mapped to a named accountable officer and a documented threshold. Where that list does not exist, direct the CRO to produce a gap list within thirty days. Where a system has a threshold but no aggregation rule, require one built against a rolling baseline, tested against at least one quarter of historical data before it is relied upon.
The Legacy Test
The measure of the risk officer who builds this register is not whether the eleven systems, or the hundredth system added after them, performed without incident during that officer's tenure. It is whether the register itself, the named owners, the specific thresholds, the working aggregation rules, is still standing and still catching drift for the successor who inherits the seat, governing whichever agentic systems are running by then, most of which have not been built yet.
The officer who builds this architecture now is acting from conviction: not built in response to litigation, but in the quiet period before an enforcement action or a derivative suit forces the question, while the doctrine is still being tested rather than settled against this exact fact pattern. The officer who waits until a regulator's sample or a plaintiff's expert finds the aggregate drift first is not building a register. They are producing an exhibit.
This three-layer boundary between board risk appetite and officer-level risk management, and the aggregation rule that catches drift a static threshold cannot see, is developed in the Executive Leadership Playbook on Agentic AI Governance, where the same standard is applied, function by function, to every officer whose name McDonald's has put on the duty of oversight.
Glenn E. Daniels II, Touch Stone Publishers