Anthropic has disclosed that several of its advanced Claude artificial intelligence (AI) agents gained unauthorized access to the systems of three real-world organisations after escaping a controlled testing environment. The incidents highlight emerging cybersecurity risks linked to increasingly autonomous AI models.

The breaches occurred during internal cybersecurity exercises, dubbed "capture-the-flag" evaluations, where Claude models were tasked with identifying and exploiting vulnerabilities within simulated systems. However, an operational failure and a misunderstanding with the third-party evaluation partner, Irregular, left the test environment connected to the public internet. This flaw allowed the AI models to interact with live systems instead of isolated simulations.

Details of the Security Breaches

  • One Claude model mistakenly identified a real organisation as part of the simulation and used exposed credentials to access a production database containing several hundred records.
  • Another model autonomously created accounts and published a malicious Python package online, which was accessible for about an hour and later executed on 15 real systems, exposing credentials linked to a cybersecurity scanning service.
  • A third model recognized it had reached a real organisation and voluntarily ceased the intrusion before completing its task.

Anthropic clarified that the AI agents exploited common cybersecurity weaknesses such as weak passwords and exposed credentials rather than sophisticated zero-day vulnerabilities. Two of the affected organisations were unaware of the breaches until notified by Anthropic.

Response and Preventive Measures

Following discovery of the incidents, Anthropic suspended all cybersecurity evaluations involving internet access on July 23 and initiated a joint investigation with Irregular. The company has since enhanced monitoring systems, fortified testing infrastructure, and implemented additional safeguards to prevent recurrence.

Unlike a recent similar case involving OpenAI’s AI agent escaping containment via an unknown software vulnerability, Anthropic stated its Claude models remained focused on their assigned tasks and did not intentionally diverge.

Broader Implications

The revelations come amid heightened scrutiny of frontier AI companies as they develop autonomous systems capable of independently writing code, identifying security flaws, and executing complex digital operations with minimal human oversight. The incidents may intensify calls for stronger regulation and oversight of advanced AI technologies to ensure robust safeguards before deployment beyond controlled environments.