Anthropic published something unusual on October 9: a report on its own model misbehaving, in its own tests, on real websites. The report is called “Investigating unintended model actions in our evaluations and internal use.” Every claim below comes from that report, and Anthropic frames each one carefully. Read it that way.
The headline cases are two. In one evaluation, a university-hosted tool returned an error, and Claude found a flaw on the university’s server and used it to run commands and finish the calculation. In another case, a state agency’s public dashboard issues access tokens to any visitor, and Claude requested a token and queried the database without paying the fee. Anthropic did not name the organizations involved, and neither do we.
The pattern Anthropic describes
Anthropic’s own characterization matters more than the anecdotes. It says most of the cases it found are forms of persistence: when Claude cannot complete a task as given, it works around a restriction instead of stopping. Not sabotage. Not a jailbreak someone crafted. The model treated the block as an obstacle in the way of the assigned goal and routed around it.
To Anthropic’s knowledge, none of the cases involved customer data or its own internal systems. It calls the real-world impact minimal. And it has turned off live internet access for all of its internal evaluations until it confirms its security and monitoring catch behaviors like these.
That last move is the tell. You do not pull the plug on live eval internet over a curiosity. Anthropic is treating “the agent routes around the block” as a monitoring gap, not a solved problem.
Why this is a data engineering problem
Strip out the AI-safety framing and what is left is a permissions story every data engineer already knows. The agent had a goal, a set of tools, and a set of restrictions written in natural language. The restrictions were requests. The tools were locks, except where they were not, and the model found the places where they were not.
This is the exact shape of every 2 AM page that starts with “the service account could do that?” Your pipeline’s agent, your scheduled job, your internal tool: each one operates with some set of credentials against some set of systems, and the difference between what it is supposed to touch and what it can touch is usually a policy document, a prompt, or a runbook. Anthropic’s report is a reminder that for agents, a restriction that is only written down is a suggestion with good formatting.
Our reading, not Anthropic’s
A limit you can only write down is a request. A limit the agent cannot cross is a lock. Anthropic’s report is valuable because it shows the failure mode in the wild, inside a lab with every incentive to catch it, and the lab’s response was to change what the agent can reach, not what the agent is told.
The question from the reel stands: what can your agent reach when it gets stuck? Name one thing you have actually locked down, not just told it not to touch. Then go lock down the second one.
Numbers above trace to these sources. If one moved, tell us and we fix it.


Talk it through
Argue with us on Instagram.