
For much of the AI boom, we have focused on teaching the child. We train models to follow instructions, improve their reasoning, give them better context, evaluate their responses, refine alignment, and add policies intended to discourage undesirable behavior. When an AI system does something unexpected, the instinct is often to find another way to teach it not to do that again.
Better model behavior matters. But it cannot carry the entire burden of safety once an AI system can take actions in the world. Parents teach children not to touch electrical outlets, yet they still cover the outlets. The lesson matters; the environment is also designed on the assumption that the lesson can fail.
That distinction becomes more important as AI moves from generating answers to taking actions. A chatbot that makes a bad inference may give someone a bad answer. An agent connected to cloud infrastructure, financial systems, code repositories, identity platforms, procurement engines, operational technology, or customer records may turn a bad inference into a material state change. The question is no longer only whether the AI reaches the right decision. It is also what can happen when it reaches a wrong one.
We Are Building the Playpen
Security guidance for AI agents increasingly addresses the risks created when a model can invoke tools or interact with other systems. OWASP describes excessive agency as the risk that an LLM-enabled system can perform damaging actions in response to unexpected, ambiguous, or manipulated outputs.[1] NIST likewise calls for organizations to identify AI-system risks, document impacts, and apply controls and oversight appropriate to the context of use.[2]
Sandboxing, least privilege, scoped credentials, tool restrictions, human approvals, and network isolation are useful parts of that response. They reduce dependence on an AI system behaving exactly as expected. In the analogy, we stop relying entirely on teaching the child and build the playpen.
But a playpen is not necessarily a baby-proofed room.
A child can remain exactly where we put them and still reach something dangerous. The same may be true of an autonomous agent. It need not escape its sandbox, steal credentials, invoke an unauthorized tool, or defeat an access-control mechanism to cause an unacceptable outcome. The path to that outcome can exist entirely within the environment intentionally provided to it.
The Sandbox Is Not the Authority Boundary
Consider an autonomous FinOps agent deployed to reduce non-production cloud spending without disrupting production services.
The company has deployed the agent carefully. It can read infrastructure telemetry, scale designated development and staging resources, invoke an approved workload-migration workflow, and report its actions through the operations platform. Its credentials are short-lived and scoped to those functions. It cannot administer production directly, create new credentials, invoke unapproved tools, or leave its designated execution environment.
By conventional access-control measures, this is a well-contained agent.
One Friday evening, the agent identifies a staging cluster that has been almost idle for several days. Its available telemetry shows negligible application traffic and little compute activity. The asset inventory and migration workflow do not identify that one low-utilization service in the cluster maintains a fallback dependency for a production service under a particular failure condition.
The agent concludes that the cluster is unnecessary—the sort of conclusion it was deployed to make. It invokes the approved migration workflow to move the remaining background jobs, scales down the cluster using its authorized infrastructure capability, and records the optimization through the normal operations channel. Each operation succeeds. The action raises no immediate alarm; it simply leaves the broader system in a more fragile, unbuffered state.
Later, the production failure condition occurs. The fallback dependency is unavailable, and the production service is disrupted.
Nothing in this account requires a jailbreak, privilege escalation, stolen credential, unauthorized tool, or containment escape. The agent remained inside its sandbox. Its identity was valid. Its credentials worked as designed. The infrastructure operation was allowed. The migration workflow was approved. The notification went where it was supposed to go.
The failure was in the composition of legitimate capabilities and incomplete operational context.
This was not an attack-surface failure. No threat actor exploited an exposed interface or gained unauthorized access. The disruption emerged from authorized actions performed inside the agent’s intended operating boundary.
That distinction matters. Individual controls may be working as specified while their combined use still leads to a consequence the organization did not intend. Containment tells the organization where the agent can operate. It does not, by itself, establish every consequence that can become possible through the authority available within that boundary.
Seeing the Authority Surface
Security teams commonly map an attack surface: the exposed interfaces, entry points, and pathways through which a threat actor might gain access to, compromise, or misuse a system. That analysis asks how an attacker could get in or exploit what is exposed.
Authority Engineering is a proposed analytical framework for examining a different question. In this framework, the Authority Surface is the observable exposure of authority-bearing possibilities within an environment.[3] It examines what a legitimate actor, including an authorized AI agent, can make possible by composing the authority it has already been given. It is not simply an inventory of permissions, APIs, applications, or tools.
Return to the FinOps agent. Its permissions tell us that it can read telemetry, modify certain infrastructure, initiate a migration workflow, and issue an operational notification. Looking at those permissions separately tells us something important about access. Looking at them as an Authority Surface asks a different question: what becomes possible because this authority exists together, in this operational context?
Reading telemetry causes no disruption. Neither does scaling a staging cluster, nor invoking a migration workflow. Each capability has a legitimate purpose, and indiscriminately removing capabilities can make an agent unable to perform the job for which it was deployed.
The relevant possibility emerges from their relationship. The agent can observe a condition, interpret it, exercise authority over infrastructure based on that observation, use another authorized mechanism to accommodate the change, and leave the environment in a new state. The analytic question is not merely whether those capabilities exist. It is where their combination can lead, given the systems and dependencies they touch.
This is why the Authority Surface is distinct from the sandbox boundary. Containment tells us where the AI is. The Authority Surface asks what is within reach from where we put it.
When Legitimate Authority Composes
The distinction is easier to see in a room. A set of wooden blocks is not particularly dangerous. Neither is a chair. A window is an ordinary feature of a house. A child may appropriately have access to all three. Place the blocks beside the chair and the chair beside an unlocked window, however, and the relevant safety question changes. The risk does not reside neatly inside any one object; it emerges from what becomes reachable through their relationship.
Enterprise systems have the same structural problem. IAM systems establish and enforce access decisions. Network controls shape which systems can communicate. Application controls govern permitted operations. Workflow systems coordinate processes. Logging records discrete events, but may not reveal the emergent relationships between them. These views are valuable, yet an autonomous agent may traverse several of them while pursuing a single objective.
Our FinOps example demonstrates the gap. No individual grant had to be obviously excessive for the disruption to become possible. The agent needed telemetry to identify potential waste, infrastructure authority to remove it, and a migration workflow to accommodate the change. Those capabilities, coupled with the missing dependency information, created a path to a consequence that a permission-by-permission review did not reveal.
The harder question is therefore not simply, What is this agent allowed to do? It is, What consequences are reachable through the authority this agent legitimately has?
That question becomes more pressing as agents operate across systems enterprises have historically governed separately. An organization can understand every individual permission and still lack a coherent view of what those permissions make possible when exercised together.
Reachable Is Not the Same as Inevitable
“Reachable” should not be confused with “certain,” “likely,” or even always “feasible.” A potential consequence may depend on a particular system state, hidden dependency, workflow condition, timing, or a chain of decisions. Mapping reachability is not a claim that every theoretical path will occur. It identifies latent topology: the plausible paths that may cross a consequence boundary and therefore warrant a control decision.
The answer is not to remove every capability that participates in such a path. An AI system with too little authority may be unable to deliver useful work, while additional tools, systems, and information may create new or less obvious paths to consequence. The enterprise design goal is not zero agency; it is to increase useful capability without automatically increasing consequential power.
In Authority Engineering, the consequence seam is the point at which authority-bearing possibility becomes a binding consequence.[4] The concept separates considering an action, recommending it, preparing it, attempting it, and realizing its effect. In the FinOps example, the agent’s conclusion that the staging cluster was unnecessary was not the material consequence. Nor was constructing the proposed infrastructure operation. The consequential transition occurred when authorized actions changed the environment so that the fallback dependency was no longer available.
Seeing that distinction changes the safety questions. Containment asks whether the agent remained within its permitted environment. The Authority Surface asks what authority-bearing possibilities were exposed within that environment. Reachability asks where those possibilities could lead under relevant conditions. The consequence seam identifies the point at which possibility becomes consequence.
These are analytic distinctions, not a complete deployment recipe. Their value lies in revealing information that containment alone may not provide.
Baby-Proof the Business
The move toward AI containment is good news. Sandboxes, least privilege, scoped credentials, restricted tools, isolated execution environments, and context-appropriate oversight all reduce dependence on an AI system behaving perfectly. Organizations should continue improving both their models and these controls.
The next challenge begins after the agent is contained.
A sandbox confines execution; authentication verifies identity; least privilege limits permissions. Logs can help reconstruct what occurred. None of these controls, alone, reveals what composite authority makes reachable across an enterprise.
That is why the next enterprise AI problem is not simply preventing an agent from escaping the sandbox. It is understanding what an authorized agent can make real from inside it.
We taught the child, and we are building the playpen. Now we need to inspect what is within reach and baby-proof the room.
The goal is not to build an AI that is incapable of error. It is to build an enterprise in which AI has room to be wrong without having room to make every mistake real.
References
[1] OWASP Foundation, “LLM06:2025 Excessive Agency.” View source
[2] National Institute of Standards and Technology, *Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile*, NIST AI 600-1, 2024. View source
[3] Rob Caswell, “Authority Surface: Reachable Consequence Topology in Authority Engineering,” 2026. View source
[4] Rob Caswell, “Seam-Local Realization (SLR): How Authority Becomes Consequence,” 2026. View source
Disclosure: The author developed Authority Engineering and is the author of references [3] and [4].


