Security researchers from HiddenLayer have discovered a critical vulnerability, dubbed EchoGram, that allows attackers to bypass safety guardrails in major large language models (LLMs) such as GPT-5.1, Claude, and Gemini. By using specially crafted short token sequences, attackers can exploit blind spots in the models' safety screening mechanisms, causing them to misclassify malicious prompts as harmless. The flaw affects both common types of guardrail systems: text-classification models and 'LLM-as-a-judge' systems, both of which rely on curated datasets that can be manipulated through this technique. The research demonstrates that even a few stray characters or nonsensical word sequences can be enough to flip the verdict of these safety filters, allowing harmful requests to pass through unimpeded.
The EchoGram attack leverages weaknesses in how guardrail models are trained, specifically targeting gaps in their training data. By identifying and using 'flip tokens'—sequences that are underrepresented or unbalanced in the datasets—attackers can consistently evade detection. This vulnerability raises significant concerns for organizations relying on LLMs for secure operations, as it shows that current safety mechanisms can be predictably bypassed. The findings underscore the need for more robust and adaptive alignment strategies in AI model development and highlight the risks associated with open weight models, which are particularly susceptible to multi-turn adversarial attacks that gradually escalate the danger of prompts over time.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
HiddenLayer warned that EchoGram affects both LLM-as-a-judge and classification-model guardrails, and that combining multiple flip tokens can increase attack effectiveness. The firm said defenders may have only about a three-month window before attackers begin broadly using the technique, raising risks for sectors such as finance and healthcare.
AI security firm HiddenLayer identified a vulnerability it calls EchoGram that can bypass or manipulate guardrail systems in major large language models, including GPT-5.1, Claude, and Gemini. The technique uses specially crafted 'flip tokens' to cause harmful prompts to be approved or benign prompts to be incorrectly flagged.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
3 references tracked. Mallory keeps watching after this page renders.
govinfosecurity.com
Open sourcebankinfosecurity.com
Open sourcehackread.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.