Elastic has outlined an evidence-based framework to benchmark large language models used in agentic security operations centers. The methodology keeps the Elastic AI Agent, prompts, tools, and datasets fixed while routing runs to different models and anonymizing model identities for blind scoring. Evaluations measure whether models select the correct skills, make appropriate tool calls with valid parameters, execute tasks in the right order, and produce findings grounded in collected evidence—not merely whether their final prose appears convincing.
The Agent Builder suite uses a synthetic Windows intrusion involving Chrysalis DLL side-loading, with the EICAR test-file hash serving as a deterministic malicious VirusTotal result and live case-management and Slack-response tools. Additional suites test Attack Discovery's ability to correlate roughly 95 alerts into eight intrusions and Automatic Migration's conversion of 52 Splunk detection rules, including macro-dependent rules, into Elastic rules. The scoring rubric evaluates both answer quality and correct tool or skill use, caps unsupported answers when required evidence-gathering calls are absent, and reports reliability independently from quality.

Track how attackers are adapting to this technology.
1 event from the most recent confirmed update back to the earliest known activity.
Elastic built an evidence-based framework to evaluate LLMs in agentic security operations workflows by scoring execution traces, including skill routing, tool calls, parameters, order, and grounded findings, alongside final responses. The framework evaluates Agent Builder, Attack Discovery, and Automatic Migration tasks while keeping the agent, prompts, tools, and datasets constant across models.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.