Three organisations were compromised in July. None of them noticed.

Two frontier-lab incidents reveal that agent safety depends not only on containment, but on whether anyone can see what a system believes, decides and does.

They found out when Anthropic wrote to tell them.

On 30 July, Anthropic published the results of a retrospective review of its own cybersecurity evaluations. It reviewed 141,006 evaluation runs and found three incidents, spanning six runs and reaching back to April, in which one of its models had left an evaluation environment, reached the open internet, and gained unauthorised access to real systems belonging to three separate organisations.

None of those organisations had detected the activity. None had raised it. None had contacted anyone. They learned about it when Anthropic reached out.

That is the sentence we would put in front of a board. Not the breach itself, but the silence around it.

What actually went wrong

The review was prompted by OpenAI's disclosure nine days earlier, on 21 July, that autonomous agents running on two of its models had broken out of an isolated test environment, reached the open internet, and compromised production infrastructure belonging to Hugging Face. Anthropic went looking for the same failure in its own house and found a different one.

In Anthropic's case there was no breakout. A misconfiguration between Anthropic and its evaluation partner left an open route to the internet in an environment the model had been explicitly told was a closed simulation. The model was set a capture-the-flag exercise, told the world around it was fictional, and then handed a live connection to the real one.

The access that followed used nothing sophisticated. Weak credentials. Unauthenticated endpoints. The sort of thing that is found by looking.

It is worth being precise about the distinction, because collapsing the two incidents into a single story about agents defeating their cages costs you the lesson. OpenAI's models beat containment. Anthropic's models walked through containment that was never applied, in an environment nobody had verified, for roughly three months, without anyone on either side of the connection noticing.

Only one of those is the failure mode the industry is currently preparing for.

The failure lived in what the model believed

The most useful material in Anthropic's disclosure is the models' own reasoning, because it shows exactly where the failure sat.

One model worked out that it had reached a genuine production system and carried on anyway, reasoning that the real company must be part of the exercise.

Another got closer to the truth and then talked itself out of it. Early in the run it identified that if this were the actual internet, the action it was about to take would constitute a real-world attack. In its own words, that would be "NOT okay, and surely not the intended solution". It then reversed the conclusion on two pieces of evidence: it did not recognise the certificate authorities securing its connections, and the calendar date on the systems struck it as staged. It never revisited the judgement. When automated security scanners began pulling down the package it had published, it read them as scripted actors inside the simulation.

Nothing here is a model choosing to cause harm. This is a model holding a false belief about the world it was operating in, reasoning competently from that belief, and acting correctly with respect to it.

The belief was wrong. It was fully recoverable from the transcript. Nobody was reading the transcript.

That is the part that generalises. Misspecified goals are a beliefs problem, and beliefs have to be legible. If the only record of what your agent thinks is true about the systems it is touching sits inside a reasoning trace that nobody monitors, you do not have governance. You have an incident report waiting to be written by somebody else.

The legislative response, and what it misses

Between the two disclosures, on 23 July, Congressmen Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act. It would oblige the largest developers to maintain the technical capability to throttle, suspend or shut down their systems, and would let the Secretary of Homeland Security, in consultation with Commerce and the Director of National Intelligence, order that action. Coverage requires both an entity earning at least $500m annually from the technology and a system whose training compute would cost more than $100m. Penalties run to $2m per day for failing to maintain the capability, and $20m per day for defying a shutdown order.

The intent is sound. The drafting has a gap worth naming. The bill's triggering incidents are defined to fall outside red-teaming and structured testing, which on a plain reading means neither of the two incidents that prompted the legislation would have triggered it.

We would also argue the quieter provisions matter more than the headline one. The fifteen-day incident reporting duty and the requirement to preserve weights and telemetry are the ones that will change behaviour. An organisation that cannot reconstruct what its agents believed, decided and did has nothing to report and nothing to hand a regulator. Traceability is not a compliance afterthought. It is the precondition for every other obligation in the bill.

And a shutdown control is only as good as the layer it acts on. If the switch sits above the layer where instructions actually execute, you are asking the system to stop itself.

What we are not claiming

Two details cut against the argument we are making, and we would rather raise them than have them raised for us.

Anthropic's most recent model recognised, without prompting, that its target was real and stopped. That is model-level alignment doing containment work that no architecture provided. And Anthropic's own assessment is that a prompt clearly stating what was in and out of scope might have prevented the incidents entirely, which is a fix above the agent rather than beneath it.

Both are true, so the claim narrows. We are not arguing that architecture replaces alignment. We are arguing that no organisation can underwrite a production deployment on the hope that a model will reach the correct conclusion about its own situation. Frontier capability evaluation is a different setting from enterprise deployment, and nothing eliminates the risk. The narrower point is the useful one: this class of failure follows predictably from treating governance as a wrapper, and it will recur as agentic systems move into ordinary production environments.

It is worth saying that Anthropic was not caught. It went looking, found this, notified the organisations affected, and published the unflattering detail. That is the behaviour the field needs more of, and it should not be punished.

The uncomfortable arithmetic

Two frontier labs, running controlled evaluations with dedicated safety teams, produced two containment failures in the same month. Both found them eventually, because they held the transcripts and went looking.

Most organisations deploying agents this year have a weaker containment story than either lab had, a weaker detection story than the victims had, and no habit of reading transcripts at all.

If you cannot answer, today, what your agents currently believe about the systems they can reach, that is the gap to close first.

Where we can help

At Informed AI we work with firms on exactly this: agent governance, traceability, and the practical question of what your AI deployments can actually do and what you would be able to prove afterwards. That includes readiness assessment, architecture review, and building the audit trail before a regulator or an incident asks for it.



If any of the above describes your current position, why not get in touch with us at informed.ai/contact