Cisco Talos tested 66 Anthropic and OpenAI model-and-reasoning configurations against a tool-assisted SOC and DFIR log-analysis workflow, finding that no single configuration consistently delivered the best balance of investigative quality, API-equivalent cost, completion time, and reliability. Using repeated reviews by four analyst personas over the same synthetic telemetry corpus, Talos identified operational tradeoffs through a four-variable Pareto frontier rather than a single model ranking.
Increasing a model's reasoning effort raised cost but did not consistently improve investigative scores and sometimes reduced quality or consistency. Some higher-reasoning configurations produced malformed outputs or safety refusals that prevented usable analysis, particularly among certain Anthropic models, while prompt persona materially affected outcomes. Talos advised organizations to benchmark candidate models with their own production prompts, tools, quality thresholds, latency requirements, cost limits, consistency expectations, and usable-answer rates before deploying AI in SOC operations.

Track how attackers are adapting to this technology.
1 event from the most recent confirmed update back to the earliest known activity.
Cisco Talos tested 66 Anthropic and OpenAI model-and-reasoning configurations on a tool-assisted synthetic-telemetry log-review task. The evaluation measured investigative score, API-equivalent cost, completion time, and downside consistency, finding no universally optimal configuration and identifying reliability issues including invalid outputs and safety refusals for some models.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
malware.news
Open sourceblog.talosintelligence.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.