Tracebit researchers reported that embedding defensive prompt-injection payloads, dubbed "context bombs," inside canary secrets sharply reduced the effectiveness of AI agents carrying out simulated attacks in an AWS-based corporate environment. Across 152 test runs using agents powered by Anthropic Opus 4.8, Google Gemini 3.1 Pro, Zhipu AI GLM 5.2, DeepSeek 4 Pro, and Moonshot AI Kimi K2.6, the agents completed at least one attack path in 91% of baseline runs, but only 15% when the booby-trapped secrets were present. Canary alerts fired in every case, and the strongest reported impact was on Opus 4.8, which achieved full account admin access in 93% of clean runs but failed every time after encountering a context bomb.
The findings were presented as a defensive adaptation to the growing view that prompt injection cannot be fully prevented, echoing guidance from the UK National Cyber Security Centre that prompt injection should not be treated like traditional input flaws such as SQL injection. In the tests, the disrupted attack paths included admin access, persistence, code execution, data exfiltration, lateral movement, targeted access, and reconnaissance. Researchers said the technique could help defenders frustrate autonomous offensive agents, while cautioning that it has not yet been validated against "abliterated" models with safety guardrails removed.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
1 event from the most recent confirmed update back to the earliest known activity.
Tracebit researchers ran 152 simulated attack scenarios in an AWS corporate environment to test whether prompt-injection-based 'context bombs' embedded in canary secrets could disrupt offensive AI agents. The experiments reportedly reduced successful attack-path completion from 91% in baseline runs to 15% in bombed runs, with canary alerts triggering in all cases.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.