SentinelLabs introduced a multi-stage benchmark to test whether frontier AI models can autonomously sustain a full malware reverse-engineering investigation, using the 2005 Fast16 Windows sabotage toolkit as the case study. Fast16 was designed to interfere with LS-DYNA engineering software and has been compared to Stuxnet because of its apparent link to efforts affecting Iran’s nuclear weapons development. The benchmark recreated an eight-stage investigation with staged evidence release, IDA Pro-based artifact generation, and verification tied to specific addresses and recovered components, including Fast16’s Lua framework and its 101-rule patching engine.
In the reported results, GPT-5.6 Sol was the only publicly available model tested that completed all eight stages, while GPT-5.5, GLM-5.2, and Opus 4.7/4.8 demonstrated useful point-in-time analysis but failed to maintain project-wide coherence. SentinelLabs said the decisive capability was project-scale recovery: retracting invalidated conclusions, tracing dependencies, and repairing the broader investigative narrative as new evidence emerged. The researchers concluded that current leading models can materially assist malware analysts and, in the strongest cases, operate as supervised investigative agents, but human reverse engineers remain necessary because the models still make semantic, technical, and quality-control errors and can declare completion prematurely.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
4 events from the most recent confirmed update back to the earliest known activity.
SentinelLabs concluded that current top AI models can materially assist malware analysis and, in the strongest cases, act as supervised investigative agents. However, the researchers said human reverse engineers remain essential because the models still make technical, semantic, and quality-control errors.
In SentinelLabs' testing, GPT-5.6 Sol was the only publicly available model reported to complete the full eight-stage benchmark. Other tested models, including GPT-5.5, GLM-5.2, and Opus 4.x variants, showed useful local analysis but failed to achieve project-wide closure.
SentinelLabs described a multi-stage, long-horizon reverse-engineering benchmark for frontier AI models by recreating its investigation of the Fast16 malware. The benchmark tested whether models could sustain a trustworthy investigation as new evidence invalidated earlier conclusions.
The benchmark centered on Fast16, a Windows sabotage toolkit/malware sample from 2005 designed to interfere with LS-DYNA engineering software. The reporting says it appears tied to Iran's nuclear weapons development context and has been compared to Stuxnet.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
2 references tracked. Mallory keeps watching after this page renders.
securityweek.com
Open sourcemalware.news
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.