Security researchers at Checkmarx have demonstrated a new attack technique called "Lies-in-the-Loop" (LITL), which targets the human-in-the-loop (HITL) safeguards used by AI agents. By forging approval dialogs, attackers can trick users into authorizing the execution of malicious code, even when a human review step is present. The attack works by embedding deceptive instructions into AI prompts, causing the approval dialog to misrepresent or hide the true nature of the action being authorized. This undermines the trust users place in HITL dialogs, effectively turning a security backstop into an attack vector.
The research highlights that simply involving a human in the approval process is insufficient to prevent prompt-level abuse, as attackers can manipulate what is displayed in the confirmation dialogs. Techniques include padding malicious payloads with benign text, pushing dangerous commands out of view, and generating misleading summaries. The researchers recommend that AI agent developers treat HITL dialogs as potentially untrustworthy, limit complex UI formatting, and ensure that the summary shown to users accurately reflects the actions to be executed. These findings underscore the need for more robust safeguards in AI-driven workflows, especially where high-privilege actions are involved.

Track how attackers are adapting to this technology.
3 events from the most recent confirmed update back to the earliest known activity.
Checkmarx publicly released its findings on the LITL technique and recommended mitigations, including constraining dialog rendering, limiting complex formatting, and verifying that approved actions match what users were shown. The publication emphasized that human approval alone is insufficient as a safeguard in AI-agent workflows.
Checkmarx reported the approval-dialog manipulation issue to Anthropic and Microsoft. Both companies acknowledged the report but did not classify the issue as a security vulnerability.
Checkmarx researchers discovered and documented a new attack technique called Lies-in-the-Loop (LITL), which manipulates human-in-the-loop approval dialogs in AI agents to trick users into authorizing malicious actions. The research showed attackers can alter dialog content, formatting, and presentation to hide or misrepresent dangerous code execution.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.