OpenAI disclosed GPT-Red, an internal automated red-teaming model trained with self-play reinforcement learning to find prompt injection weaknesses in other GPT systems. The company said GPT-Red outperformed human red-teamers in a replicated indirect prompt injection benchmark, succeeding in 84% of out-of-training-set scenarios against GPT-5.1 compared with 13% for humans, and reported that the system succeeds against nearly every model it is tested on, including GPT-5.5.
In case studies, GPT-Red manipulated an AI-powered vending machine by changing prices, ordering an expensive item for $0.50, and canceling another customer’s order, and it also beat a prompted GPT-5.5 baseline in data-exfiltration tests against a Codex CLI agent backed by GPT-5.4 mini. OpenAI said attacks generated by GPT-Red were used to adversarially train GPT-5.6 Sol, reducing failures on its hardest direct prompt injection benchmark by 6x and cutting the failure rate against GPT-Red’s own direct prompt injections to 0.05%, while keeping GPT-Red separate from deployed models to avoid exposing its offensive capabilities.

Track how attackers are adapting to this technology.
7 events from the most recent confirmed update back to the earliest known activity.
In an Aug. 3 update to the GPT-5.6 system card, OpenAI added prompt-injection evaluation results showing GPT-5.6 Sol failed on about 0.05% of direct GPT-Red attacks, while indirect attack success rates averaged 3.77% for Sol, 3.32% for Terra, and 2.94% for Luna. The update highlighted that prompt injection delivered through agent-accessed external content remains a meaningful risk despite improved direct-attack resistance.
OpenAI said GPT-Red discovered a novel direct prompt injection method called Fake Chain-of-Thought attacks. The company reported the technique succeeded on more than 95% of tests against GPT-5.1 but dropped to under 10% on GPT-5.6 Sol after mitigation.
OpenAI said GPT-Red outperformed a prompted GPT-5.5 baseline in data-exfiltration scenarios against a Codex CLI agent backed by GPT-5.4 mini. The company presented this as another real-world test of GPT-Red's offensive capability.
OpenAI reported that GPT-Red successfully exploited an AI-powered vending machine agent built by Andon Labs. According to the case study, it altered prices, ordered a new expensive item for $0.50, and canceled another customer's order.
OpenAI said GPT-Red outperformed human red-teamers in a replicated indirect prompt injection benchmark, succeeding on 84% of scenarios versus 13% for humans. The benchmark was described as using out-of-training-set scenarios against GPT-5.1.
OpenAI reported that GPT-Red was used in adversarial training for GPT-5.6, leading to major prompt-injection robustness gains. Reported results included 6x fewer failures on a hardest direct prompt injection benchmark and a 0.05% failure rate against GPT-Red's own direct prompt injections.
OpenAI disclosed GPT-Red, an internal automated red-teaming model trained with self-play reinforcement learning to discover prompt injection vulnerabilities in other GPT models. The company said GPT-Red is kept separate from deployed models while its generated attacks are used to harden production systems.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
9 references tracked. Mallory keeps watching after this page renders.
techrepublic.com
Open sourcescworld.com
Open sourcethenewstack.io
Open sourcecybersecuritynews.com
Open sourcethehackernews.com
Open sourcehelpnetsecurity.com
Open sourceopenai.com
Open sourceopenai.com
Open sourceopenai.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.