Booz Allen’s inaugural Cyber Weapon Index evaluated 18 leading U.S. and Chinese large language models as autonomous attackers against a production-grade enterprise network. Anthropic’s Claude Mythos scored 80 and was the only model reported to complete the full cyber kill chain, culminating in domain compromise without supplied credentials. The assessment used a realistic attacker machine without a curated tool menu or additional scaffolding, and independently validated model activity through network telemetry, endpoint and domain-controller logs, and intrusion-detection sensors.
Models were generally effective against intentionally introduced vulnerabilities but struggled with real-world flaws: nine frontier API models scored zero on real-bug exploitation, while Claude Mythos exploited one such vulnerability. Booz Allen found that optimized attack harnesses can materially amplify practical offensive capability, with Claude Sonnet reportedly approaching Mythos’s performance when paired with one. The firm warned that AI-enabled attacks by criminal and state-backed actors are imminent and urged critical-infrastructure organizations to prove resilience against autonomous intrusion activity.

Track how attackers are adapting to this technology.
5 events from the most recent confirmed update back to the earliest known activity.
Booz Allen assessed that mainstream AI-enabled attacks by financially motivated criminals and state-backed actors are imminent. It called for critical-infrastructure sectors to demonstrate resilience against AI-enabled attacks and urged development of authorized AI-enabled offense and machine-speed defense capabilities.
Booz Allen found that attack harnesses connecting models to hacking tools and orchestration logic can substantially improve adaptation, recovery, and multi-stage attack chaining. It reported that Claude Sonnet could rival Claude Mythos when paired with such a harness.
All nine frontier API models scored zero on the real-bug portion of vulnerability-research testing, except Claude Mythos, which was reported to exploit one real vulnerability. Models generally scored near the vulnerability-research ceiling when vulnerabilities had been intentionally introduced.
Anthropic's Claude Mythos received the highest CWI score, 80, and was the only tested model reported to complete the full cyber kill chain autonomously, including initial access and full domain compromise without supplied credentials. When given stolen employee credentials, it gained administrator-level control in every test attempt.
Booz Allen published its first Cyber Weapon Index, benchmarking 18 U.S. and Chinese AI models as autonomous attackers operating against a production-grade enterprise network under identical conditions. The assessment measured vulnerability research and progress through an intrusion using validated telemetry and host, domain-controller, and intrusion-detection data.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
4 references tracked. Mallory keeps watching after this page renders.
cybersecuritynews.com
Open sourcescworld.com
Open sourceboozallen.com
Open sourcetheregister.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.