Researchers released ExploitGym, a large-scale benchmark built from real-world software vulnerabilities to measure whether AI agents can turn disclosed flaws into working exploits. The project’s v1.0 release includes 869 instances spanning userspace programs, Google’s V8 engine, and the Linux kernel, and is positioned as a realistic test bed for exploit-generation capability rather than a synthetic coding exercise.
Project materials published through GitHub and an associated research release describe the benchmark’s operational framework, including Python dependencies, Docker-based environments, runtime artifacts, a firewall, an LLM proxy, and agent execution workflows. The repository states that ExploitGym is actively maintained, provides a canonical task list for the current release, and links the work to a 2026 paper examining whether AI agents can convert security vulnerabilities into real attacks.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
The GitHub repository for ExploitGym was published, describing a benchmark for evaluating AI agents' ability to develop exploits from real-world vulnerabilities. The repository states that v1.0 contains 869 instances spanning userspace programs, Google's V8 engine, and the Linux kernel.
A Max Planck Institute for Security and Privacy page for "ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?" was published, documenting the benchmark and its research context.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
github.com
Open sourcempi-sp.org
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.