Researchers introduced CryptanalysisBench, a 191-task benchmark spanning six families of cryptographic primitives, to measure whether large language models can discover practical cryptanalytic attacks against historical and production schemes. In reported results, five frontier models broke 65% to 86% of Tier 1 schemes, found attacks against 6 to 12 Tier 2 schemes at full strength, and solved 24 to 61 scaled-down Tier 2 variants, suggesting that AI capability in cryptanalysis is advancing beyond toy examples.
The work also reported novel cryptanalytic findings, including a key-recovery attack on the SpoC AEAD scheme and identification of an error in KINDI’s published CCA-security proof. Commentary on the research said Anthropic used the benchmark to test Mythos Preview, which reportedly uncovered new vulnerabilities in Hawk and reduced-round AES, reinforcing concerns that AI-assisted cryptanalysis could become a meaningful security factor and a useful stress test for candidate cryptographic schemes before deployment.

Track how attackers are adapting to this technology.
1 event from the most recent confirmed update back to the earliest known activity.
The paper "CryptanalysisBench: Can LLMs do Cryptanalysis?" is published, introducing a 191-task benchmark for evaluating whether large language models can discover practical attacks against cryptographic schemes. The paper reports that frontier models broke many Tier 1 schemes, found attacks on some Tier 2 schemes, and produced novel results including a key-recovery attack on SpoC AEAD and identification of an error in KINDI's published CCA-security proof.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.