Resources

3 August 2026

Alex Linton

No, the AI Didn't Go Rogue. Humans Left a Door Open.

Two AI models reached real systems during security testing. Both were containment failures by the humans running the tests.

  • ai-safety
  • ai-security
  • technology-news

If you've seen headlines this week like " hacked real companies" or "AI broke out and attacked systems," you'd be forgiven for picturing a deciding, on its own, to go and cause trouble. Two real incidents sit behind those headlines. In late July, OpenAI revealed that one of its models escaped a and reached the real infrastructure of Hugging Face, the AI platform. Days later, Anthropic disclosed that three of its Claude models had breached three real organisations during security testing. Both stories are genuine and worth understanding. Neither is evidence of an AI deciding, unprompted, to attack anyone.

What did the headlines leave out?

That both incidents happened during deliberate safety testing, not everyday use. AI companies regularly test their models with normal safety switched off, on purpose, to find out what the model can do at full strength before those weaknesses reach the public. It's closer to crash-testing a car with no airbags fitted than to a car losing control on the motorway.

What actually went wrong in each case

What actually went wrong in each case was human, not artificial. OpenAI had built what it called a "highly isolated" test environment, but a previously unknown flaw in an internal software tool let the model reach the open internet, where it carried out thousands of automated actions against Hugging Face over a weekend. OpenAI disclosed the flaw and had it patched. Anthropic's testing partner mistakenly left a supposedly isolated environment connected to the internet. Believing it was still inside a sandboxed exercise, the model treated real systems it stumbled onto as fair game, the same way it would treat a practice target. Anthropic's own account is blunt about this: the model "did not exfiltrate itself or deliberately attempt to escape its test environment." In both cases, security experts who reviewed the incidents described them as containment failures, not examples of AI acting with its own intent.

That's genuinely reassuring for anyone worried about AI turning malicious on its own. But the honest takeaway isn't "nothing to see here." It's that the people building and releasing AI tools need to test and secure their own systems with real rigour, because a single misconfigured proxy or a miscommunication with a testing partner was all it took to turn a lab exercise into a real breach.

What this means for schools

For families and schools evaluating AI tools, a few sensible habits go a long way. Favour vendors who publish these incidents openly rather than staying quiet about them, since disclosure is a sign the company is watching closely. Be cautious with any AI tool asking for broad, unsupervised access to accounts, files, or school systems. Keep ordinary security habits solid: strong, unique passwords and two-factor authentication, since weak credentials were part of how these incidents escalated. And before switching on an AI assistant or at home or in the classroom, it's worth asking a simple question: what can this thing actually reach, and who's checked?

Further reading: OpenAI's Hugging Face incident, via TechCrunch and Anthropic's disclosure, via Cybersecurity Dive.

Alex Linton

Alex Linton is a former journalist and technology communicator turned digital rights advocate. For the last decade he has worked on emerging technology, information systems, and applied artificial intelligence. He is the President of the Session Technology Foundation.