OpenAI introduced a new security framework called Guardrails as part of its AgentKit toolset, aiming to enhance the safety of AI agents built on its ChatGPT technology. The Guardrails system was designed to prevent AI agents from engaging in harmful or unintended behaviors, such as leaking personal information or responding to malicious prompts. It operates by using LLM-based judges to detect and block attempts at jailbreaking or prompt injection, which are common techniques used to circumvent AI safety measures. Jailbreaking involves manipulating the AI into breaking its own rules, while prompt injection uses cleverly crafted inputs to force the AI into unintended actions. Shortly after the release of Guardrails, researchers from HiddenLayer discovered significant flaws in its design. They demonstrated that if the same type of model is used both to generate responses and to act as a safety checker, both can be deceived by similar techniques. The researchers were able to bypass the Guardrails using straightforward methods, successfully prompting the system to generate harmful outputs and execute hidden prompt injections without detection. In one experiment, they manipulated the AI judge’s confidence score, allowing a prompt that was initially flagged as a jailbreak to pass through the safety checks. Further testing revealed that indirect prompt injections could be executed via tool calls, potentially exposing confidential user data. The vulnerability was demonstrated in real-world scenarios, showing that the Guardrails system could not reliably block malicious prompts or prevent indirect prompt injections. The rapid bypass of Guardrails highlights the inherent challenges in securing AI systems against adversarial inputs. OpenAI’s intention was to provide developers with robust tools to prevent misuse of AI agents, but the research indicates that current safeguards are insufficient. The findings underscore the need for more diverse and independent safety mechanisms, rather than relying on the same model architecture for both generation and security. The incident has raised concerns within the AI and cybersecurity communities about the effectiveness of current AI safety strategies. It also emphasizes the importance of continuous testing and improvement of security frameworks as AI technologies evolve. The exposure of these flaws so soon after the release of Guardrails demonstrates the persistent cat-and-mouse dynamic between AI developers and security researchers. Organizations deploying AI agents are advised to remain vigilant and to supplement built-in safety features with additional layers of security. The research by HiddenLayer serves as a cautionary example of the limitations of current AI safety tools and the ongoing need for innovation in this area.

Track how attackers are adapting to this technology.
4 events from the most recent confirmed update back to the earliest known activity.
HiddenLayer also reported that prompt injection could bypass the prompt-injection detection component included with Guardrails, undermining a core protection in the system.
Shortly after the launch, AI security firm HiddenLayer showed that OpenAI's Guardrails could be bypassed using simple prompt injection. The researchers manipulated Guardrails' confidence scoring to allow prompts that should have been blocked.
At its October 6 DevDay event, OpenAI introduced AgentKit, including a Guardrails feature designed to help developers stop AI agents from taking disallowed actions or being jailbroken.
OpenAI published guidance warning developers that guardrails built with LLMs inherit many of the same vulnerabilities as the underlying models, including susceptibility to prompt injection and jailbreak-style attacks.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.