New benchmark results indicate that the framework surrounding an AI model is often the decisive factor in both cybersecurity and desktop automation tasks. On macOS, the MacAgentBench benchmark measured 676 tasks across 25 applications and found that Claude Opus 4.6 paired with the OpenClaw framework reached 73.7% first-try success, while the same model using only screenshot-based mouse-and-keyboard control achieved 39.2%; GPT-5.4 led bare-agent setups at 58.4%. Researchers said much of OpenClaw’s advantage disappeared on tasks without prebuilt recipes, suggesting that skill libraries and task-specific scaffolding, rather than model quality alone, accounted for most of the gains. Reliability also remained uneven: the best setup completed 85.2% of tasks at least once in four attempts but only 58.6% correctly on all four runs.
Cybersecurity-focused evaluations reported a similar pattern. Comparative reviews of CyberMetric, NYU CTF Bench, InterCode CTF, and CAIBench found that models perform well on knowledge-heavy tests but drop sharply on multi-step exploitation and attack-defense tasks, where success rates often fall into the 20–40% range. A separate CyBench study discussed on r/netsec found no single scaffold dominated every challenge, while combining heterogeneous scaffolds and a shared blackboard architecture solved 19 of 33 tasks (57.6%) and reduced execution time. Safety findings also pointed to the harness as the main control point: one SkillHarness study said explicit skill boundaries and risk guards enforced user-consent and sensitive-data restrictions before execution, and that removing those boundaries increased attack success by 9.6 percentage points. Across the studies, researchers warned that stronger automation without robust guardrails could enable unauthorized file access, credential harvesting, and other misuse.

Track how attackers are adapting to this technology.
4 events from the most recent confirmed update back to the earliest known activity.
A write-up on SkillHarness reported that safety is enforced through explicit skill boundaries and planner-checked risk guards tied to environment state, such as requiring user consent or blocking sensitive-data submission. It cited an ablation study showing that removing skill boundaries increased attack success by 9.6 percentage points, supporting the conclusion that the harness boundary provides the primary safety function.
A 2026 arXiv paper discussed on r/netsec reported benchmarking five cybersecurity scaffolds using the same alias2-mini model across 33 CyBench challenges. The study found no single scaffold was best on every challenge, while a shared blackboard architecture solved 19 challenges, or 57.6%, and reduced execution time.
An analysis comparing CyberMetric, NYU CTF Bench, and InterCode CTF reported that LLMs perform strongly on knowledge-based cybersecurity tests but much worse on multi-step offensive tasks such as exploitation and attack-defense scenarios. It cited CAIBench as showing a persistent gap, with knowledge benchmarks above 70% success and multi-step attack scenarios often only 20–40%, while also noting that scaffolding can affect outcomes more than model choice.
Researchers reported results from MacAgentBench, a benchmark of 676 tasks across 25 macOS applications, finding that Claude Opus 4.6 with OpenClaw achieved 73.7% first-try success while bare-agent setups performed substantially worse. The reported analysis concluded that much of the gain came from framework-provided skill libraries and task recipes rather than the underlying model alone, and also highlighted reliability and misuse concerns.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
4 references tracked. Mallory keeps watching after this page renders.
reddit.com
Open sourcemedium.com
Open sourcehelpnetsecurity.com
Open sourcecodeby.net
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.