What Directors Should Be Able to Demonstrate Before an AI System Fails
A Touch Stone Publishers White Paper for Boards of Directors
The central governance problem created by artificial intelligence is not that directors must become model engineers. It is that a board can approve material AI use without establishing a reliable way to learn when the system has moved outside the conditions the board believed it was approving.
That is the fiduciary vacuum. Management can demonstrate technical performance, the vendor can describe safeguards, and counsel can explain the legal framework, yet the board can remain unable to produce a contemporaneous record showing what information reached it, which signals required escalation, who owned the response, and what changed after a warning appeared.
This paper advances a narrow claim. For a material AI system, the board's most defensible contribution is a documented oversight architecture that existed before the failure and operated while the relevant decisions were being made. The architecture does not guarantee that the system will perform correctly. It gives the board a disciplined way to know when performance, conduct, or dependency risk requires attention.
Touch Stone Publishers calls that architecture the Documented Oversight Loop. It consists of four connected records: a reporting cadence, an owned threshold registry, a vendor responsibility map, and a contemporaneous response log. Together, they convert oversight from a statement of intent into an observable management system.
Scope note: This paper presents a governance design, not legal advice. No Delaware court had decided an AI-specific Caremark claim as of August 31, 2026. The application of existing oversight doctrine to AI is therefore reasoned analysis, not a prediction of liability in any specific case.

The Board's Problem Is Evidentiary Before It Is Technical
Boards often begin the AI discussion with the performance of the technology. They ask whether the model is accurate, whether the vendor is credible, whether a pilot produced useful results, and whether management has established controls. Those questions are necessary. They are not sufficient.
The board's distinct question is whether it has established a reasonable information pathway for a risk that may be central to the enterprise. Existing Delaware oversight doctrine focuses on good-faith efforts to establish and monitor information and reporting systems. Cases such as Marchand v. Barnhill and the Boeing oversight litigation illustrate the importance of board-level visibility into mission-critical risk. They do not decide an AI case, and they do not turn every business failure into director liability. They do make one principle difficult to ignore: a board cannot monitor information it has not arranged to receive.
AI makes that principle operationally demanding. The technology is embedded in workflows, supplied through layered vendors, updated frequently, and capable of changing the speed and scale of decisions. A quarterly statement that management has observed "no material issues" can be sincere and still be too weak to show how a material warning would reach the board. The problem is not necessarily bad faith. It is an information design that is too general for the risk it is supposed to surface.
The distinction matters because technical assurance and fiduciary oversight answer different questions. Technical teams ask whether a system performs within an approved operating envelope. The board asks whether the organization has a repeatable method for bringing material departures, dependencies, and unresolved warnings into the governance process. A strong technical program can support that method. It cannot substitute for it.
The same distinction appears in emerging regulation. The FDA's August 2026 publication on generative AI-enabled medical devices is a discussion paper seeking feedback, not a final policy change. Its attention to risk assessment, competency evaluation, and postmarket monitoring nevertheless shows the type of lifecycle information that may become important in regulated uses. The European Union's AI Act establishes documented risk-management and governance obligations for defined categories of high-risk systems. Executive Order 14420 addresses specified foreign-produced bulk-power equipment and related supply-chain dependencies, not AI systems as a blanket category. These instruments differ in scope and legal effect. Their common governance implication is narrower: boards need an information system capable of distinguishing what applies to a particular use case, what has changed, and who must respond.

Why General Risk Oversight Is Not Enough
The strongest objection to a dedicated AI oversight architecture is organizational, not doctrinal. General counsel, the audit committee, the risk committee, internal audit, and the technology function already have responsibilities for legal compliance and enterprise risk. Adding another structure can create duplication, dilute accountability, and encourage governance theater.
That objection is correct about the danger of creating a new committee for every emerging issue. It does not follow that AI can be absorbed into a generic technology-risk paragraph.
A general risk process works only when it retains enough specificity to surface the condition that matters. If a material AI system affects a regulated asset, a customer decision, a safety process, or a critical operational dependency, the reporting system must identify that system, its escalation conditions, its accountable owner, and its vendor dependencies. A general statement about "AI risk" does not tell a director what would trigger an urgent report or whether anyone is monitoring that trigger.
The answer is not a separate AI bureaucracy. It is specificity inside the governance structures that already exist. A risk or audit committee can retain ownership. General counsel can coordinate the legal analysis. Management can remain responsible for operations. The improvement is that material AI systems appear by name in the board information flow, with defined thresholds and owners. That design strengthens the existing structure instead of creating a parallel one.
A second objection is more serious. If the board focuses on producing artifacts, the organization may satisfy the form of oversight while leaving the substance unchanged. A named agenda item, a register, or a response template can become paperwork that no one uses.
The Documented Oversight Loop addresses that risk by connecting each record to a management action. The cadence requires someone to gather and present current information. The threshold registry requires operational leaders to define which conditions matter before a warning appears. The vendor map requires procurement, legal, risk, and technology leaders to reconcile their assumptions. The response log requires a decision and follow-through to be recorded while it occurs. The artifact is not the objective. It is evidence that the operating discipline exists.

An Illustrative Decision That Exposes the Vacuum
Consider a composite healthcare services company preparing to renew and expand a generative AI clinical-documentation platform. The original system drafts notes that clinicians review. The proposed expansion adds preliminary triage suggestions, bringing the use closer to clinical judgment and increasing the consequences of undetected degradation.
Management presents strong adoption and satisfaction results. The vendor reports no material incidents. Two directors ask who would detect a decline in performance and how the board would learn about it. Management explains that the vendor performs quality reviews and will report problems. The board approves the expansion. The minutes state that the vendor relationship was discussed and approved.
Several weeks later, a clinician notices that the system is recommending lower acuity for a particular complaint. The concern reaches a department head informally. The department head raises it during a routine vendor meeting. The vendor describes the pattern as expected variance and continues internal monitoring. The organization has no defined internal threshold for the signal, no named escalation owner, and no reporting pathway independent of the vendor whose performance is in question.
The governance weakness does not depend on predicting the clinical outcome. It is visible in the information architecture. The board approved a material expansion without a record of the conditions that would trigger escalation, the internal owner responsible for monitoring them, the vendor responsibility gap, or the pathway by which the board would learn that a warning had appeared.
The same approval could have been rational under a stronger system. A board does not demonstrate good governance by rejecting every AI initiative. It demonstrates discipline by making the basis of the decision, the accepted residual risk, and the monitoring conditions explicit. Under the Documented Oversight Loop, the record would show the reporting cadence, the system-specific thresholds, the current vendor responsibility map, and the response protocol. The board could accept the expansion while preserving the ability to see and act on material information.

The Documented Oversight Loop
The Documented Oversight Loop is deliberately narrow. It does not certify that an AI system is safe, accurate, fair, secure, or compliant. Those conclusions require technical, operational, legal, and independent assurance appropriate to the use case. The loop governs the board-facing evidence that those functions are working, that important warnings move upward, and that responses move back into operations.
Record One: A Documented Reporting Cadence
The reporting cadence defines what information reaches the board or responsible committee, who provides it, and how often. It names material systems rather than collapsing them into a general technology update. It distinguishes routine reporting from event-driven escalation. It also identifies the decision rights attached to the report, so that the board knows whether it is receiving information, approving a risk decision, or confirming a remediation plan.
The corporate secretary or general counsel can maintain the governance record. Management owners provide the underlying measures and analysis. The board does not need every model metric. It needs a reliable view of the measures connected to the approved use, the material risk, and the escalation conditions.
A cadence is credible when it survives personnel changes and quiet periods. If reporting occurs only because one director asks unusually good questions, the organization has not established a system. It has established a dependency on that director.
Record Two: An Owned Red-Flag Threshold Registry
The threshold registry translates technical, operational, compliance, and vendor signals into defined escalation conditions. A threshold may involve a failed assessment, an unexplained performance change, a customer or employee impact, a vendor breach notice, a regulatory inquiry, a material change in data or model dependency, or an inability to explain a consequential output.
Each threshold requires a named owner, an expected response, and a time boundary. This does not mean every signal reaches the full board. It means the pathway is designed before the signal appears. Some thresholds belong with a management risk committee. Others require the audit or risk committee. A small number should trigger full-board notification.
The registry is strongest when operating teams participate in its design. General counsel can frame the governance requirement, but the people closest to the system understand how failure first appears. The board's role is to ensure that the thresholds are material, owned, and periodically tested.
Record Three: A Scrutinized Vendor Responsibility Map
Material AI systems often depend on several external parties: foundation-model providers, application vendors, data sources, implementation partners, infrastructure providers, and monitoring services. Contract language, operational control, and practical response capacity may sit with different parties.
The vendor responsibility map makes those dependencies visible. It compares who operates each control, who receives each warning, what remedies are available, and where the organization retains exposure even when the vendor caused the underlying problem. It does not declare the legal result. It identifies the questions that procurement, counsel, the risk function, and qualified advisers must resolve.
The map matters because a signed agreement can create false confidence. Procurement may confirm that contractual protections exist while the board assumes that material risk has been transferred. The governance task is to see the difference between contractual recourse and the organization's continuing operational, regulatory, financial, and reputational exposure.
Record Four: A Contemporaneous Response Log
The response log records what the organization knew, when it knew it, who reviewed the signal, what decision was made, and what follow-through was assigned. It is created as the response unfolds, not reconstructed after the outcome is known.
The board or committee record should remain at the governance level. Management should maintain the operational chronology. The two must connect. The board record shows that the right information arrived, that the appropriate decision-maker engaged, and that follow-through was monitored. The operational record shows the investigation, containment, correction, validation, and return-to-service decisions.
The log closes the loop only when the response changes the system. A threshold may need revision. Reporting may need to become more frequent. A vendor control may need renegotiation. A use case may need to be narrowed or paused. A response that leaves the next reporting cycle unchanged is an event record, not a learning system.

The Board's Place in the Operating Model
The board's role is to establish the information architecture and hold the management system accountable. It should not perform technical validation, manage incidents, or negotiate routine vendor terms. The distinction is essential because governance becomes weaker when directors either receive too little information or drift into management's work.
The practical division of responsibility is straightforward. The board or responsible committee approves the oversight design and material risk decisions. The corporate secretary and general counsel maintain the board record and ensure that defined matters reach the agenda. Management owners monitor the systems and thresholds. Procurement and legal maintain the vendor map. Internal audit or another qualified assurance function tests whether the process operates as described.
Independent assurance has a specific purpose in this model. It should not merely confirm that documents exist. It should select a material system and trace a sample warning from source signal through threshold assessment, escalation, decision, response, and closure. That test reveals whether the four records form a loop or remain separate files.

A 30-Day Board Action Path
A board can establish the first version of this architecture without creating a new committee or convening a special governance summit. The work fits within ordinary agenda management, pre-read preparation, risk review, and board resolutions.
| Timing | Board action | Accountable owner | Evidence produced |
|---|---|---|---|
| Days 1 to 5 | Add a recurring Material AI Systems item to the responsible committee calendar | Committee chair and corporate secretary | Updated calendar and charter language |
| Days 1 to 10 | Identify AI systems that affect regulated assets, consequential decisions, safety, or critical operations | General counsel with CIO or CTO | Named material-system inventory with executive owners |
| Days 6 to 15 | Define red-flag thresholds for the highest-exposure systems | System owners coordinated by risk and counsel | Threshold registry with owners and escalation paths |
| Days 10 to 20 | Map material vendor dependencies, controls, notice duties, and unresolved responsibility gaps | Procurement and legal | Vendor responsibility map and decision memorandum |
| Days 15 to 25 | Adopt a response-log standard and event-driven notification condition | Corporate secretary and committee chair | Approved template and minuted escalation rule |
| Days 21 to 30 | Trace one real or simulated warning through the entire loop | Internal audit or qualified assurance owner | Test result, corrective actions, and closure record |
The sequence matters. Boards often begin with a diagnostic score, identify weaknesses, and stop at the gap analysis. A defensible operating process moves from identification to ownership, from ownership to decision, and from decision to tested evidence. The day-30 output is not a maturity label. It is the result of a trace test showing where the loop operated and where it broke.
The board should resist false precision. A score can help organize discussion, but it does not prove effectiveness. The stronger question is whether the organization can produce the four records for one material system and demonstrate that they connect. Once that works, the architecture can be extended to the rest of the inventory.

A Named Escalation Condition
Every oversight architecture needs a condition that operates outside the regular calendar. Without it, a quarterly reporting cadence can become a reason to delay a material warning.
A board should adopt an event-driven notification rule tailored to the company's risk profile. One possible starting point is written notification to the responsible committee when a material system experiences a failed competency assessment, a defined performance or conduct threshold, a vendor breach notice, a regulatory inquiry, or another condition listed in the threshold registry. The rule should specify the owner, recipient, required information, and time boundary.
The appropriate time boundary depends on the event, applicable law, contractual duties, and the organization's operating context. The board should obtain qualified advice before adopting a fixed rule. The governance principle is independent of the exact number of hours or days: a material signal should move according to a pre-agreed pathway, not wait for the next scheduled meeting or depend on informal judgment about whether the board would want to know.

The Board Self-Test
The most useful test begins with one material AI system and the actual record, not a survey response.
Identify the last substantive report the board or responsible committee received about the system. Locate the current threshold that would require escalation. Name the executive who owns the response and the person responsible for the vendor side of the issue. Then produce the most recent response record and show what changed as a result.
If those answers require memory, scattered correspondence, or a narrative assembled for the exercise, the board has identified the point at which oversight depends on reconstruction. That is the place to improve first.
The test should be repeated after a simulated warning. The organization should introduce a plausible signal, route it through the threshold process, convene the correct decision-makers, document the response, and confirm that the action returns to the reporting cadence. This is not incident theater. It is a test of whether the information system the board believes it has can operate under pressure.

The Legacy Test
An oversight system must outlast both the director who championed it and the AI system that first made the weakness visible.
The first requirement is institutional continuity. Reporting cadence, escalation rules, and evidence standards belong in charters, calendars, templates, and assigned roles. They should not depend on one director's vigilance or one executive's personal knowledge.
The second requirement is architectural continuity. The framework should apply to new systems without assuming that every system carries the same thresholds or risk. The four records remain stable; their content changes with the use case. That balance allows the board to govern a changing portfolio without rebuilding its oversight model for every vendor and without forcing different risks into one generic checklist.
The legacy test therefore asks whether the organization has built a durable capacity to see, decide, and learn. A one-time AI review does not meet that standard. A documented loop that is assigned, tested, and improved can.
Successors should inherit a working oversight system, not a founder's or director's private memory. The system should be built from conviction before a crisis. It is not built in response to litigation, enforcement, or public harm.

Conclusion: Build the Record Before the Event
Boards do not need certainty about the future of AI law to improve the quality of oversight today. They need precision about their own information system.
The fiduciary vacuum appears when management operates a material AI system, vendors control important dependencies, and the board lacks a reliable way to see the signals that matter. More policy language does not close that gap. A documented operating loop does.
The reporting cadence brings material information forward. The threshold registry defines what requires escalation. The vendor responsibility map exposes dependencies and unresolved assumptions. The response log records the decision and turns the event into an improved control.
The decisive board question is not whether the organization has an AI policy. It is whether the board can produce the record showing that oversight was operating before a failure made that record necessary.

Evidence and Authority
This paper distinguishes controlling authority, regulatory material, and Touch Stone analysis. The principal source record includes Delaware oversight decisions and opinions, the FDA's August 2026 discussion paper on generative AI-enabled medical devices, the European Union AI Act, and Executive Order 14420 on specified bulk-power equipment and supply-chain risk.
- Delaware Courts, Boeing oversight opinion
- Delaware Courts, officer duty of oversight opinion
- FDA, considerations for generative AI-enabled medical devices
- European Union, consolidated Artificial Intelligence Act
- The White House, Executive Order 14420
Evidence date: August 31, 2026. Later legal, regulatory, contractual, and insurance developments should be verified before this framework is applied to a specific organization.
Author: Glenn E. Daniels II, Touch Stone Publishers