Why a Documented Decision-Rights Boundary That Is Not Enforced in Configuration Is a Policy Statement, Not a Control

Touch Stone Publishers | Glenn E. Daniels II

A documented decision-rights boundary forks into two outcomes: a policy statement that only describes intent, or an enforced control that a runtime configuration actually carries out.

The problem, stated in the CIO/CTO’s own language

The Governance Boundary Principle translated into a technical enforcement stack: board policy artifact, versioned decision-rights configuration, and runtime escalation logic.

A Chief Information Officer or Chief Technology Officer who has read a decision-rights policy for an agentic system and never asked to see the configuration that enforces it has read a document, not audited a control. This is the technical distinction this white paper exists to make. It matters whenever a board, auditor, regulator, or litigant tests whether the organization did what its policy said. In In re McDonald’s Corporation Stockholder Derivative Litigation (Del. Ch., January 25, 2023), the Delaware Court of Chancery held that corporate officers owe an oversight duty within their areas of responsibility. The court also emphasized that liability requires bad faith, not an imperfect system or an ordinary mistake.

Other officers build the organizational record: a named accountable person, a documented boundary, and a review cadence. The CIO/CTO builds the technical substrate that makes those artifacts testable. A board minute naming an accountable officer is stronger when the logging architecture can show what that officer was notified of, when the notice occurred, and what action followed.

Stated as a single technical-architecture proposition: a decision-rights boundary that exists in a policy document but is not encoded as an enforced constraint in the agent’s runtime configuration is not an operating control. It is a statement of intent. The gap between what a policy says an agent may decide and what the system will actually permit is a gap only the CIO/CTO’s organization can close.

This is an architecture problem with governance and evidentiary consequences. The 30-day sequence in this paper establishes a minimum viable control for one priority system. It does not complete an enterprise governance program.

Why this is a discovery-evidence problem before it is a compliance problem

What a policy document proves versus what an enforced configuration proves: stated intent and a named role on one side, a boundary actually respected and testable against the running system on the other.

Caremark doctrine distinguishes a bad-faith failure of oversight from an imperfect system or an ordinary mistake. Red flags can matter, but they are not the entire legal test. For a technology officer, the practical question is whether the information system was reasonably designed to surface material risk within the officer’s area of responsibility.

Agentic systems break that assumption at the architectural level, not the policy level. An agent can execute thousands of individually unremarkable decisions that aggregate into a material harm with no single decision looking, in isolation, like a flag worth raising to anyone. There is no moment at which a human, however diligent, would have been expected to notice, because the harm is a property of the aggregate, not any individual transaction. This is the finding the underlying Playbook calls the black box problem, and it means the statement “no one raised it with me” provides little operating assurance. The system may never have been designed to identify and route the aggregate pattern.

The management response is architectural: a demonstrated, working, documented oversight system that does not depend solely on a human noticing one unusual decision. It should surface defined risk mechanically, on a cadence or trigger, to a named accountable person. That system exists in the logging schema, retention schedule, boundary configuration, and alerting logic, not only in a policy binder.

If litigation, regulatory review, internal audit, or a board inquiry tests the system, the organization may need prompt records, model and configuration history, decision logs, and architecture diagrams. An organization that never created or retained these artifacts may be unable to demonstrate what its oversight process actually did. This paper does not prescribe a universal retention period. Legal counsel, records management, privacy, security, and technical owners should determine the applicable schedule before deployment.

That asymmetry gives the CIO/CTO’s technical build direct evidentiary weight. The technical record does not replace governance documents. It makes their claims testable.

The Governance Boundary Principle, translated into runtime enforcement

The Governance Boundary Principle holds that the board governs and management manages, and that an organization begins to fail, quietly at first and then suddenly, the moment either crosses into the other’s territory. The underlying Playbook applies this principle to agentic systems through a three-layer architecture: the board sets category-level policy for what may ever be delegated to an agent without a human sign-off; a named accountable officer holds documented authority over a specific system’s configuration within that boundary; and the agent executes only what has been granted to it inside that boundary, with a defined mechanism for surfacing behavior that approaches or crosses it.

For every other function in the underlying Playbook, that three-layer architecture is organizational: a charter amendment, a named title, a signed document. For the CIO/CTO, the same three layers describe a technical build, and the Governance Boundary Principle’s warning about boundary violations applies with equal force in code as it does in a boardroom.

When engineering hard-codes a decision threshold into an agent’s runtime without the board’s category-level policy ever having authorized that category of delegation, engineering has crossed into the board’s territory, the same failure mode the Governance Boundary Principle names in any other operational context, except that here the violation is buried in a configuration file instead of a memo, which makes it harder to discover and more dangerous once discovered. When a board or an outside director starts asking to review individual model configurations, prompt templates, or threshold values directly, rather than the category-level policy those technical choices should sit inside, the board has crossed into management’s territory in a domain where technical judgment, not board judgment, is the competency required, and where that crossing slows exactly the adoption speed Section 0 of the underlying Playbook establishes as separately, severely costly.

The CIO/CTO’s job under this framework is precise: build the system so that the board’s policy boundary, once set, is the boundary the system actually enforces, not a boundary the system happens to respect until someone changes a configuration value without anyone else knowing. That requirement has three concrete technical corollaries, each one a control a court, a regulator, or an internal auditor can test directly against the running system rather than against a document describing it.

First, the decision-rights boundary approved by the board or the named accountable officer must be represented as a version-controlled artifact that the agent’s runtime actually reads and enforces, not a separate document that engineering references informally when building the system and then does not revisit. Second, any change to that boundary, whether a threshold, a category of authorized decision, or an escalation condition, must itself be logged, attributed to a specific change author, and time-stamped, so that the boundary’s own history is as reconstructable as the agent’s decisions inside it. Third, the system must be built so that a decision outside the documented boundary is either technically impossible or automatically flagged the moment it occurs, not merely against policy in a way that depends on a subsequent human audit to catch.

An organization that has done this work has converted the Governance Boundary Principle from an organizational value the board asserts into an enforced technical constraint a regulator can inspect. An organization that has not done this work has a policy document that describes a boundary the running system may or may not actually respect, and neither the CIO/CTO nor anyone else in the organization can say which, with confidence, until something goes wrong and the configuration is finally examined under pressure.

Function-specific risk quantification: what the CIO/CTO personally carries

The underlying Playbook is explicit that the CIO/CTO’s direct exposure runs through the technical evidentiary trail that discovery in this domain now targets specifically: prompt logs, model version history, and architecture diagrams. This is a distinct and personal exposure, separate from the general technology risk a CIO/CTO has always carried, for three reasons.

First, it is retroactive in a way most technology risk is not. A security vulnerability discovered today can be patched today. A missing log of a decision made six months ago cannot be recreated after the fact; the evidence either exists in retained form or it does not, and the retention decision that determined the answer was made, and owned, well before anyone knew it would matter. The CIO/CTO who sets a routine operational log rotation policy for an agentic system without a cross-functional retention review has made a decision today that cannot be undone in the future, no matter how much technical capability the organization has by then.

Second, it is a decision no other officer in the organization is positioned to catch. The General Counsel can review a decision-rights document for completeness. The CFO can review a certification for accuracy against a stated control. Neither can inspect a logging configuration and know, without the CIO/CTO’s own technical judgment, whether it actually retains what a future future inquiry will require. This makes the CIO/CTO the technical custodian of the evidence that supports every other officer’s paper trail, not merely a peer contributor to the organization’s overall governance posture.

Third, regulated sectors already impose recordkeeping, model-risk, fair-lending, privacy, and consumer-protection obligations that can affect the design. Those obligations vary by jurisdiction, entity, use case, and record type. NIST’s voluntary AI Risk Management Framework offers a broader governance reference: document roles, maintain inventories, measure performance, monitor production behavior, and track risk over time.

The career-risk consequence for the CIO/CTO personally is specific and structural rather than reputational in the way a public failure is for a CEO or a COO. A technology leader whose systems cannot reconstruct what an agent decided and why, after the fact, has left the one gap every other officer’s governance record depends on the technology function having closed. When a General Counsel’s decision-rights document, a CFO’s certification, or a Board Committee’s minute is tested in discovery, all three rest on the CIO/CTO’s logging and retention architecture having actually captured what those other documents assert happened. An officer whose own function is sound but whose organization’s technical substrate cannot support it inherits a failure they did not cause and cannot, after the fact, correct.

The worked scenario: underwriting logging, before and after, at full technical depth

The underwriting logging architecture before and after: schema fields, retention window, and escalation trigger logic rebuilt so reconstruction is a query, not an investigation.

The following underwriting scenario is illustrative. It is not a description of a named case or a universal technical prescription. It shows how logging architecture, retention policy, and escalation logic change the quality of the evidence an organization can produce.

Before: a system that recorded outcomes but did not retain the evidence needed to reconstruct how it reached them.

An agentic underwriting system logged its final decisions: approve, decline, refer for manual review, along with the applicant record and the decision outcome. It did not preserve a structured decision basis, the relevant retrieved factors, the specific model version and configuration active at the moment of any given decision, or the approved prompt-template version. The logging schema had been designed by the engineering team responsible for the underwriting pipeline’s operational reliability, not by anyone tasked with anticipating a future discovery request, and it was optimized accordingly: enough to debug a system outage or reconcile a daily batch, nothing more.

Retention followed the same operational logic. Decision logs were retained for ninety days, the standard window the platform team applied to every service in the environment for troubleshooting purposes, governed by a generic infrastructure retention policy that had never been reviewed by legal, privacy, security, records-management, and business owners. No one had connected the underwriting system’s log retention to the questions the organization might later need to answer. The retention clock ran on a rolling basis tied to storage cost management, not evidentiary need.

Six months after a cluster of adverse underwriting decisions began drawing internal complaints, a compliance review asked engineering to reconstruct what had happened: which model version had been active, what inputs had been considered, and why a specific category of applicant had been declined at a materially higher rate than a comparable category. Engineering could not answer any of these questions with confidence. The model had been updated twice in the intervening period, without a corresponding change log tied to decision records. A structured decision basis had never been captured at all. The original decision logs, to the extent they existed beyond the ninety-day window, had rotated out of retention months before the review began. What remained was the final decision and the applicant record, exactly the two data points least useful for answering the question the review was actually asking: not what the system decided, but why, and under whose configuration.

This is the discovery-readiness failure the Playbook identifies. The missing record does not by itself establish bad faith or liability. It does prevent the organization from answering basic questions about which system version acted, what basis was recorded, who reviewed the result, and whether an alert fired.

After: a system built so the reconstruction is a query, not an investigation.

The rebuild addressed three distinct technical layers, each mapped to a specific requirement the underlying Playbook’s Section 4 diagnostic names under Discovery Readiness and Escalation Mechanism.

Logging architecture. The underwriting system’s logging schema was rebuilt to capture, for every individual decision, four categories of record rather than one: the final decision and applicant record, as before; the specific model version and configuration hash active at the moment of the decision; the structured decision basis, such as retrieved factors, a scoring breakdown, or an approved outcome code, sufficient to reconstruct the material basis without retaining private model reasoning; and the identity and authority of any human who touched the decision at any stage, including a reviewer who approved a referral. This is a structured, append-only event log tied to a unique decision identifier, not a set of scattered application logs that happen to mention the decision in passing. Each decision produces one retrievable record, queryable by applicant, by date range, by model version, or by outcome category, rather than requiring an engineer to reconstruct the picture from multiple disconnected systems after the fact.

Retention policy design. Retention for this class of record was deliberately separated from the organization’s generic operational log retention policy. Legal counsel, records management, privacy, security, and the system owner approved a record-specific schedule tied to the applicable business and legal requirements. The policy was written as an explicit, named exception inside the platform’s broader data lifecycle management, with its own owner, so that a future cost-optimization initiative could not silently roll the record back into the default operational window. The retention policy document itself, not only the underlying data, was version-controlled and reviewed on the same cadence as the decision-rights boundary it was built to support, so that a reviewer auditing the architecture could confirm the retention commitment matched the retention behavior actually configured in the system, rather than trusting that the two had stayed aligned by default.

Automated escalation triggers. The technical escalation mechanism was rebuilt to trigger automatically on defined behavioral thresholds rather than depending on a human periodically querying the logs, which is precisely the distinction Section 1 of the underlying Playbook identifies between passive red-flag monitoring and an architectural control. The trigger logic monitors the decision stream for two categories of signal: a single-decision threshold, where any individual decision crosses a defined severity or dollar exposure line and is routed to a named reviewer before, not after, it takes effect; and an aggregate drift threshold, where the approval or decline rate for a defined applicant category deviates beyond a stated percentage from a trailing baseline period, which is the mechanism built to catch exactly the failure pattern the original incident represented, a harm invisible at the level of any single decision but visible immediately at the level of pattern drift. When either trigger fires, the system generates a notification to the named accountable officer automatically, with a direct link to the specific decision records that produced the alert, and logs the fact that the notification was sent and when, so the escalation mechanism’s own operation becomes part of the reconstructable record, not merely the underlying agent’s decisions.

The combined effect changes the quality of the organization’s evidence. Under the before state, a compliance review or a discovery request asking what happened and why produced an investigation that could not be completed, because the evidence required to answer the question no longer existed. Under the after state, the same question produces a query: retrieve the decision identifier, retrieve the associated model version, structured decision basis, and any escalation events tied to it, and produce an answer inside hours rather than failing to produce one after months. That is the difference between an architecture that can be tested and one that cannot be reconstructed. The difference was built at the level of logging schema, retention policy, and escalation logic, not at the level of a revised policy document.

What the CIO/CTO owes the rest of the organization, on a fixed timeline

The CIO/CTO's 30-day action plan mapped against the Section 4 Governance Readiness Diagnostic bands, moving the most exposed system one full band up.

The underlying Playbook recommends a thirty-day foundation for moving its most exposed agentic system at least one full band up the Section 4 Governance Readiness Diagnostic. For the CIO/CTO, that standard resolves into four concrete deliverables, each one testable against the running system rather than against a stated intention.

The first week’s work is an honest audit: every mission-critical agentic system scored specifically against the diagnostic’s Discovery Readiness dimension, producing a gap list rather than a general impression of the organization’s maturity. The following week targets the highest-priority gap identified in that audit, implementing retained logging sufficient to reconstruct an individual decision, with the retention window approved through the cross-functional process described above rather than inherited from the shorter default used for operational troubleshooting. The test of success is specific: a reconstruction of a real past decision, attempted using only the retained logs, either succeeds or it does not. The third week builds or upgrades the automated escalation mechanism itself, so that it fires on a defined threshold without depending on a human’s initiative to go looking, tested against a simulated breach rather than assumed to work because it was built with good intentions. The final week produces a single technical architecture memo, written at a level of specificity the General Counsel can incorporate directly into the officer-level decision-rights documentation built across the rest of the organization, and the test of that memo’s sufficiency is not the CIO/CTO’s own judgment. It is the General Counsel’s confirmation that the memo actually supports the organization’s documented oversight claim in a form that would survive the kind of scrutiny this white paper has described throughout.

That final handoff is the point of the exercise. Every other officer’s governance record is written in the language of accountability, authority, and documented boundaries. The CIO/CTO’s technical build makes those claims testable against the running system.

The two-clause test this architecture has to pass

The two-clause Legacy Test applied to a technical architecture: what survives the officer's tenure, and whether it was built from conviction before an incident or as an exhibit after one.

The measure of the CIO/CTO who builds this system is not whether the agentic systems under their watch happened to perform well while they held the role. Performance alone is not the measure of an oversight architecture. The Legacy Test asks two architectural questions.

First, what the documented, working oversight architecture, the logging schema, the retention policy, the escalation logic, leaves behind for the technology leader who inherits the seat, governing whatever agentic system is running by then, on whatever model, under whatever vendor relationship has by that point replaced the one in place today. An architecture built to survive a specific system’s replacement, rather than merely describe the current system, gives the successor a usable control environment. The duty described in McDonald’s attaches to the officer’s area of responsibility and is evaluated under a demanding bad-faith standard. This technical design is Touch Stone’s recommended evidence architecture, not a judicially mandated specification.

Second, whether this architecture was built from conviction, in the quiet period before an enforcement action or a derivative suit forced the question, or only afterward, in response to one. An organization that builds retained logging, aligned retention policy, and automated escalation triggers before any incident has occurred has built a genuine control. An organization that builds the identical architecture only after an inquiry has already found the gap has built an exhibit, not a control, and the distinction is visible to anyone who checks the dates.

This two-part Legacy Test is developed in the Executive Leadership Playbook on Agentic AI Governance. It is applied to every officer named across that work, from the board’s committee charter language through the CEO’s public statements to the technical substrate this white paper has described. The CIO/CTO who builds this architecture now, deliberately, in code, before anyone asks for it, is the leader this test was written for.


Sources and scope

  1. Delaware Court of Chancery, In re McDonald’s Corporation Stockholder Derivative Litigation, C.A. No. 2021-0324-JTL (January 25, 2023): https://law.justia.com/cases/delaware/court-of-chancery/2023/c-a-no-2021-0324-jtl.html
  2. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework 1.0 and Playbook. The framework is voluntary: https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
  3. U.S. Department of Justice, Evaluation of Corporate Compliance Programs (September 2024), including testing whether controls work in practice and managing emerging technologies: https://www.justice.gov/criminal/criminal-fraud/page/file/937501/dl
  4. NIST, AI Agent Standards Initiative (February 17, 2026). The initiative announces work toward standards and guidance; it does not create a binding technical standard: https://www.nist.gov/news-events/news/2026/02/announcing-ai-agent-standards-initiative-interoperable-and-secure

The versioned boundary, event-log fields, escalation design, diagnostic bands, and 30-day sequence are Touch Stone Publishers recommendations. Retention periods and record content require system-specific legal, privacy, security, and records-management review. This paper is governance analysis, not legal advice.

Choose your next step

Continue the collection or save the topic. Return to the Agentic AI Governance collection and choose the paper most relevant to your current decision.

Continue or save this topic

Stay current. Receive new research and executive intelligence as it is published.

Join the newsletter

Go deeper. The annual intelligence membership provides the governed research, implementation detail, and continuing updates behind the public library.

Request a membership seat