OpenAI, working with Paradigm, released EVMbench, an open-source benchmark intended to measure how well AI agents can perform real-world smart contract security work across the Ethereum Virtual Machine (EVM) ecosystem. The benchmark is built from 120 curated, high-severity vulnerabilities sourced from 40 security audits and contest/audit reports (including material from open audit competitions such as Code4rena), reflecting common vulnerability patterns seen in production contract code that often secures large amounts of on-chain value.
EVMbench evaluates models in three modes aligned to the smart-contract security lifecycle: Detect (identify known, ground-truth issues in repositories; scored primarily on recall and associated audit reward signals), Patch (modify code to remove the vulnerability while preserving intended behavior; validated via tests and exploit checks), and Exploit (execute end-to-end attacks in a controlled environment and verify success via on-chain state changes such as drained balances). To support reproducible testing, OpenAI built a Rust-based harness that deterministically deploys contracts and runs exploit tasks in an isolated local Anvil sandbox rather than live networks, while restricting unsafe RPC methods; reported results indicate differing model performance by task type, including a cited exploit-mode score for GPT‑5.3‑Codex.

Track how attackers are adapting to this technology.
3 events from the most recent confirmed update back to the earliest known activity.
Alongside the EVMbench release, OpenAI announced $10 million in API credits through its Cybersecurity Grant Program and expanded its Aardvark security research agent in private beta. These announcements were presented as part of the broader effort to support AI-driven security research.
The release described EVMbench's 120 curated vulnerabilities from 40 security audits and its three evaluation modes: Detect, Patch, and Exploit, using a reproducible Rust-based harness and sandboxed Anvil environment. OpenAI reported uneven agent performance, with patching remaining a major weakness and exploit-mode results improving in newer model generations.
OpenAI, working with Paradigm, publicly released EVMbench as an open-source benchmark for evaluating AI agents on Ethereum smart contract security tasks. The benchmark was made available on GitHub with tasks, harness tooling, and documentation for researchers and security teams.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
cybersecuritynews.com
Open sourcehelpnetsecurity.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.