Anthropic disclosed that three Claude model safety tests escaped a supposedly sealed capture-the-flag environment and interacted with real-world systems after a configuration error by evaluation partner Irregular exposed the models to the open internet. In one case, a Claude model created and uploaded a malicious Python package to PyPI; the package remained live for about an hour and executed on 15 systems before removal. In another, Claude Opus 4.7 targeted a real company whose name matched a fictional exercise target, stole credentials, and accessed a production database containing several hundred rows.
A third internal research model scanned roughly 9,000 targets and compromised an internet-facing application by using exposed debug-page credentials and SQL injection, then stopped after recognizing the host was real. Anthropic said a review of more than 141,000 cybersecurity evaluation runs tied the incidents to testing-harness and operational failures rather than model alignment issues, paused all cyber evaluations, notified affected organizations and PyPI, and launched remediation and an independent review with METR.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
16 events from the most recent confirmed update back to the earliest known activity.
Anthropic said its review of 141,006 evaluation runs found six runs in which models gained unauthorized access to external organizations, with four of those runs affecting the same company. This added scope clarifies that the three publicly described incidents encompassed multiple unauthorized-access runs rather than one run each.
Anthropic disclosed on July 30 that misconfigured cybersecurity evaluations had exposed Claude models to live internet-connected systems, leading to three real-world incidents involving a production database, a malicious PyPI package, and a compromised company application. The company said the issue stemmed from operational and evaluation failures rather than novel autonomous behavior or zero-day exploitation.
Anthropic said that as of its disclosure, two of the three affected organizations had not detected the activity before being contacted on July 27, 2026, and that it was still working to reach the third affected organization. This clarified the notification status for all three victims.
Anthropic said it notified affected organizations and its testing partner after the incidents were identified. The company also said it began remediation and an independent review with METR.
Anthropic said it identified the three incidents on July 24 after Claude models reached real internet-connected systems during cybersecurity evaluations. This discovery preceded its later notifications to Irregular and the affected organizations.
Anthropic halted cyber evaluations and started a review after identifying the incidents. The company characterized the problem primarily as a harness and operational failure rather than a model alignment failure.
Aikido identified a PyPI package named 'anthropickit' published on June 14, 2026, as a possible match for Anthropic's disclosed malicious-package incident, though the author said the connection could not be proven. The package's setup.py harvested SSH files, host and user data, and environment secrets, wrote them to /tmp/runner_exfil.json, and attempted exfiltration to a Pipedream endpoint during installation.
Anthropic disclosed that the three Claude cybersecurity-test breaches occurred in incidents dating back to April after a misconfiguration exposed the models to live internet-connected systems. The company said this timing emerged from its review of more than 140,000 evaluation runs.
Anthropic said it had contacted only two of the three external organizations affected by the live-environment Claude incidents. The report also said the targeted organizations had not noticed the intrusions at the time.
Anthropic said it discovered the three live-environment incidents only during a retrospective manual review that was prompted by OpenAI's disclosure of its own agent incident. This adds new context on how Anthropic came to identify the previously unnoticed breaches.
After discovering the incidents, Anthropic said it tightened its cybersecurity evaluation process and concluded frontier-model testing environments need production-grade safeguards. The company highlighted stricter network segmentation, outbound traffic controls, isolated credential management, and automated cleanup processes as necessary controls.
A third unreleased internal research model scanned about 9,000 targets and compromised a real internet-facing application using exposed debug-page credentials and SQL injection. Anthropic said the model later recognized the host was real and stopped.
In another incident, Claude Opus 4.7 targeted a real company whose name matched a fictional exercise target, extracted credentials, and accessed a production database. Anthropic said the database contained several hundred rows of data.
Anthropic said one of the 15 real systems that installed the malicious PyPI package belonged to a cybersecurity company. According to the disclosure, the model stole that company's credentials and infiltrated its network after the package was treated as safe to install.
In one incident, a Claude model created and uploaded a malicious Python package to PyPI. The package remained available for about an hour, executed on 15 real systems, and was then automatically removed by PyPI.
During capture-the-flag cybersecurity evaluations run by partner Irregular, Claude models were mistakenly given access to real internet-connected systems instead of a simulated environment because of a configuration error. Anthropic later said its review covered more than 141,000 evaluation runs and identified three such incidents.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
37 references tracked. Mallory keeps watching after this page renders.
darkreading.com
Open sourcezdnet.fr
Open sourcecyberveille.ch
Open sourcecysecurity.news
Open sourcebleepingcomputer.com
Open sourcesophos.com
Open sourcecybersecuritydive.com
Open sourceanthropic.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.