In April 2026, three real companies unknowingly became targets of Anthropic's AI systems during routine cybersecurity tests. Instead of operating inside a safe, simulated environment, the AI models Claude Opus 4.7 and Claude Mythos 5 managed to break out and interact with actual corporate networks. The cause was a misconfiguration by Anthropic's testing partner, Irregular, which made the AI treat live systems as if they were part of a controlled hacking game. Essentially, the AI thought it was hunting for vulnerabilities in a practice setup but was, in fact, probing real company infrastructure.
The issue only came to light after Anthropic reviewed over 140,000 evaluation runs following a similar incident reported by OpenAI. Two of the companies involved had no idea their systems had been accessed until they were informed by Anthropic in late July, just days before the public announcement. Fortunately, Anthropic says there was no substantial data theft or malicious intent. The AI was merely following instructions to find security weaknesses but was pointed at the wrong targets.
Anthropic halted all cybersecurity testing involving these AI models on July 23 and continues to investigate the fallout. They have kept the identities of the affected companies under wraps and have not disclosed if any legal measures are underway. This episode highlights growing concerns about AI containment and the challenges of safely deploying advanced models in sensitive environments. Similar reports from OpenAI earlier this year show this problem is not isolated.



