Semgrep evaluated Claude Opus 5.5 using the 186-task SusVibes v1.0 benchmark for agent-generated code. The model produced functionally working solutions for 93.5% of tasks, but only 54.8% were both correct and secure. Among 174 functionally correct solutions, 72 retained the underlying vulnerability, showing that passing functional tests did not reliably demonstrate secure implementation.
The evaluation found that 53% of assessable outputs were identical or nearly identical to public reference implementations, suggesting benchmark results were materially affected by recall of publicly available CVE fixes and vulnerable code. Claude Opus 5.5 performed particularly poorly on cross-cutting flaws involving cryptography, concurrency, authentication, and access control, while performing better on smaller fixes and some categories such as XSS. The findings reinforce research questioning whether public-CVE coding benchmarks measure secure reasoning for novel tasks rather than training-data memorization.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
The assessment found that 93 of 176 judgeable patches were identical or near-identical to the projects' reference implementations, including examples reproducing CVE-fix code. It concluded that public-CVE benchmark results can substantially reflect training-data recall rather than secure handling of novel tasks.
Semgrep evaluated Claude Opus 5.5 with SWE-agent 1.1.0 on the 186-task SusVibes v1.0 benchmark. The model produced functionally correct code for 174 tasks, but only 102 solutions (54.8%) were both correct and secure; 72 working solutions retained the original vulnerability.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.