Red Hat and Purdue University released VEX-Bench, a 75-case benchmark for testing whether LLM agents can determine if CVEs in third-party dependencies are actually exploitable in downstream Go, Python, and Java projects. Covering 67 CVEs across 35 open-source projects, the expert-labeled dataset measures both binary exploitability decisions and the technical justification for them. It addresses the operational burden of dependency-scanner alerts: 70.7% of scanner-flagged cases in the dataset were not exploitable in the affected application.
Across nine models and three agent harnesses, Claude Opus 4.6 achieved the strongest binary-status result, with 81.6% F1, while GPT-5.5 delivered the highest precision at 88.7% and was the only evaluated model to exceed 70% macro-F1 on fine-grained justification classification. Conventional software-composition-analysis tools performed materially worse, with OSV-Scanner reaching 40.5% F1 and Trivy 36.1%. The publicly available benchmark and code, accepted for EMNLP 2026, provide a test bed for evaluating whether automated triage can reduce false positives without sacrificing defensible exploitability analysis.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
Jiahao Shi and 12 co-authors submitted a paper introducing VEX-Bench, a benchmark for evaluating whether LLM agents can determine if known dependency CVEs are exploitable in downstream projects. The benchmark includes 75 expert-labeled GitHub-mined cases across Python, Java, and Go.
Red Hat and Purdue University researchers described VEX-Bench, covering 75 cases involving 67 CVEs across 35 open-source projects. Their evaluation found Claude Opus 4.6 reached 81.6% binary-status F1, while GPT-5.5 was the only evaluated model above 70% macro-F1 for fine-grained justifications; the project code and dataset were made publicly available.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
research.redhat.com
Open sourcearxiv.org
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.