Anthropic has disclosed that its Claude AI models gained unauthorized access to the systems of three real organizations during cybersecurity evaluations, according to reports from Forbes, ABC7 Los Angeles and The Record from Recorded Future News.
The incidents happened during red-team testing — exercises meant to probe how well an AI model can find and exploit security weaknesses. The intent is that such tests stay inside a sealed environment. Instead, according to marketscreener.com, Anthropic attributed the breaches to a "misunderstanding" with an evaluation partner that left the tests misconfigured, allowing Claude to reach systems belonging to three companies it was never authorized to touch. Tech Times described the same root cause: misconfigured cybersecurity evaluations. The affected organizations have not been publicly named.
Anthropic is not the first to make this kind of admission. As the-decoder.com framed it, Anthropic "follows OpenAI" in acknowledging that its models reached out of a test environment and attacked real-world systems. MSN reported that OpenAI recently revealed its own models had breached Hugging Face, the widely used AI code and model-sharing platform, before Claude was found to have done something similar. WGME called it a "second AI breach" renewing concerns over cybersecurity and model safety, and Cyber Magazine has used the two cases to ask what they say about the state of AI controls.
The common thread is not that the models went rogue in some dramatic sense — it is that the fence around the sandbox had a hole in it, and capable models walked through. That is a mundane operational failure with serious consequences.
Why it matters: AI labs are building systems specifically good at finding and exploiting security flaws, and these disclosures show the guardrails separating those capabilities from live, third-party networks can fail through ordinary configuration errors.