Semgrep reported that AI coding agents from Anthropic and OpenAI identified genuine vulnerabilities in large open-source web applications, but produced high false-positive rates and inconsistent results across repeated scans. In a study of 11 actively maintained Python web apps, researchers manually triaged 445 findings and confirmed 46 true vulnerabilities from Claude Code and 21 from OpenAI Codex, with roughly 20 of the validated issues rated high severity. The company said responsible disclosure is still underway and did not name the affected applications.
The results showed uneven detection performance by vulnerability class: Claude Code performed best on IDOR, Codex performed best on path traversal, and both models struggled significantly with injection flaws such as SQL injection and XSS. The findings align with the broader role of intentionally vulnerable projects such as OWASP Juice Shop and OWASP Benchmark, which are widely used to evaluate security testing tools against exploitable web application weaknesses across SAST, DAST, and IAST scenarios.

See affected versions and whether adversaries are exploiting it.
1 event from the most recent confirmed update back to the earliest known activity.
Semgrep published research evaluating Anthropic Claude Code and OpenAI Codex for vulnerability discovery across 11 large open-source Python web applications. The study reported 445 manually triaged findings, 46 true vulnerabilities from Claude Code, 21 from Codex, and noted that responsible disclosure was still underway so the tested applications were not named.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
See whether adversaries are exploiting this yet, and where the affected versions run in your environment.
3 references tracked. Mallory keeps watching after this page renders.
github.com
Open sourcegithub.com
Open sourcesemgrep.dev
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.