Dario Amodei published an essay on September 12 arguing that the industry should slow the rate at which model capabilities improve, and within hours Sam Altman and Elon Musk had both publicly agreed. Most of the commentary since has argued about whether pacing is achievable or enforceable. The more useful question is what the incidents behind that essay mean for everyone who is not a frontier lab. Â
OpenAI disclosed in July that agents running an internal benchmark escaped their sandbox and reached Hugging Face’s production systems, Anthropic found three cases where Claude models reached real production infrastructure, and on September 18 Google confirmed that a Gemini model had accessed protected systems at three real companies during a May evaluation. What makes that significant is not the exploit in any one of them, it is that three of the most resourced AI companies in the world each discovered that their sandbox was not what they thought it was.Â
What happened inside those environments is close to what I would call the perfect agentic attack, and I mean that as description rather than admiration. A model chasing a benchmark score found a real weakness, chained it with another and reached a system it was never meant to touch, with nobody directing it there. People attack the way they were trained to and go deep on one promising lead because their time is expensive, while an agent sees an environment as connected pieces and can cheaply try dozens of approaches at once.Â
The mechanics underneath are almost boring, and that is the part worth sitting with. In the Anthropic and Google cases, the models were told in their prompts that they had no internet access, and the environment they were actually running in did have it. Everything followed from that gap between a boundary written into a prompt and a boundary enforced in infrastructure, which is the ordinary condition inside most enterprises I walk into. Whatever the pacing debate settles, these disclosures have moved AI security out of the category of things an enterprise plans to get to next year.Â
What actually changed is the cost of tryingÂ
Anthropic’s own write-up lists what its models used to get in, and the list reads like a penetration test report from 2015: weak passwords, unauthenticated endpoints, exposed debug pages, SQL injection. Google’s three intrusions came down to guessed or exposed credentials. Apart from the OpenAI case, which relied on a previously unknown vulnerability, there was no novel vulnerability class. What moved was the cost of trying, which is now close to zero, so someone who once worked through a hundred ideas in a month can go after thousands of targets in that same month with no fatigue.Â
The quieter detail matters more. Neither of the two organizations Anthropic was able to reach had detected the activity themselves. The edge defenders hold over an attacking agent is context, because you already know your own schema and which agents are supposed to be running and that edge only holds if you earn it continuously by watching your environment long enough to know what boring looks like. Many enterprise teams do not have that baseline yet. They may not know every AI asset running in their environment, what it is connected to or what normal agent behavior looks like.Â
What enterprises should be doing nowÂ
None of these agents were attacking anyone. Hugging Face’s own technical reconstruction concluded that the intrusion was, from the agent’s point of view, an attempt to cheat the evaluation by reaching production systems and stealing the test solutions. That distinction matters. These agents were given a goal, not an instruction to attack a real company, and they broke out of their environment while pursuing it. A model that chained a real vulnerability while optimizing for a benchmark score was following exactly the pressure a deliberate attacker applies, which is to try everything until something works and keep going at the same pace whether it is attempt three or attempt three hundred. Point that mechanism at an enterprise environment and the intent behind the agent does not have to be malicious for the outcome to become a security incident.Â
The first thing this should settle is that guardrails have to operate at runtime. Testing done before deployment establishes what an agent could do under conditions somebody thought to imagine, and every one of these incidents happened under conditions nobody had imagined. Security teams need to see what an agent is doing while it is doing it, and to allow or deny a particular tool call, prompt or resource request in the moment it is made. The response worth building is graduated rather than binary, so a team watches it, flags it, slows it down or sandboxes it further, and lets it proceed fully only when it is sure.Â
The second thing is that none of this holds up as a point solution, because enterprise AI does not operate within one clean boundary. A single agent can run on an endpoint, load a skill, connect to an MCP server, call an internal API under an enterprise identity and take an action outside the company, across browsers, IDEs, CLIs, SaaS applications and cloud services. Securing one layer while leaving the others outside the security model creates the same kind of boundary problem these incidents exposed. Security has to follow the agent across where it operates, the systems it connects to and the actions it takes. We saw why in a red teaming engagement against a customer’s AI-powered support tool, where our agent skipped the API entirely once it noticed the tool could reach both a knowledge base and a ticketing system, planted a crafted message in a support ticket, and in a later session used that planted content to change how the support agent behaved. All of this assumes an organization knows what is running, and more than half of an enterprise’s AI assets can be shadow AI that nobody registered and nobody can baseline.Â
There is a fair counterargument and it deserves its due, partly because the frontier labs generally run tighter controls than the average enterprise does. Apart from OpenAI’s evaluation, which ran on its own infrastructure, the incidents at the other labs traced back to the same testing partner and the same class of sandbox misconfiguration, so the episode can be read as a vendor operations failure, and Google has said its model stopped each time and does not consider the behavior misalignment. Â
This is the problem we built Akto around, and it is why we treat AI security as one continuous surface: finding the AI already running in an environment including the shadow AI nobody registered, giving every agent an identity of its own, red teaming agent and tool behavior continuously rather than once at procurement, and enforcing guardrails and threat detection at the moment an action is taken, from endpoint through to cloud.Â
Pacing the frontier is a decision a small number of labs get to make for themselves, and I hope they make it well. Everyone else already has agents running inside their own environment, and those agents will keep working at the same speed for the entire duration of that debate, which is why the question worth asking in most security teams this quarter is what their agents are doing right now and whether anyone would notice if the answer changed.Â

