Researchers disclosed an architectural flaw in the reasoning APIs used by OpenAI, Anthropic, and Google that allowed hidden chain-of-thought or opaque reasoning objects from stronger models to be replayed into weaker compatible models and returned in plain text. The issue reportedly stemmed from reasoning envelopes not being cryptographically bound to a specific model, user, or session, enabling recovery of concealed reasoning, embedded secrets, and potentially harmful hidden content using only normal API access.
In tests across 6,708 public agent trajectories from platforms including GitHub and Hugging Face, the researchers said they decoded 315,320 thinking blocks and found 704 privacy artifacts from real user sessions, including API keys, passwords, access tokens, private keys, and other PII. The researchers warned the flaw could also support stealthy indirect prompt injection against autonomous agents, and said they disclosed the issue to the affected providers as well as Microsoft and Hugging Face; OpenAI, Anthropic, and Google acknowledged or addressed the findings with server-side mitigations that reportedly stopped the original proof-of-concept attacks.

Track how attackers are adapting to this technology.
9 events from the most recent confirmed update back to the earliest known activity.
OpenAI, Anthropic, and Google deployed server-side mitigations that made the original cross-model replay proof-of-concept attacks no longer reproducible. The paper's reproducibility statement said the main extraction attack was no longer reproducible as of August 2026.
According to the report, Matthew Green submitted the replay behavior to OpenAI and Anthropic through their bug-bounty programs. The later paper cited this as earlier related disclosure activity.
Earlier research by Johns Hopkins cryptographer Matthew Green showed that encrypted reasoning blocks could be replayed across sessions and accounts in OpenAI and Anthropic systems. This prior work established the replay behavior that later research expanded into a broader extraction attack.
A collaborative team from the ELLIS Institute Tübingen, Max Planck Institute, MATS Research, and Snyk publicly disclosed the architectural flaw affecting OpenAI, Anthropic, and Google APIs. They described how provider-wide authenticated reasoning envelopes could be replayed across users, sessions, and models to expose hidden chain-of-thought and secrets.
All three affected AI providers acknowledged the research findings about hidden reasoning exposure and replay across compatible models. Their acknowledgments were reported alongside the disclosure of the flaw.
The researchers reported the issue to OpenAI, Anthropic, Google, Microsoft, and Hugging Face. The disclosures covered the replay and extraction risks associated with hidden reasoning objects in provider APIs.
The paper included a proof of concept in which a malicious instruction was embedded inside an opaque reasoning block and replayed into an unrelated task. The receiving model carried out an attacker-directed upload action without exposing the injected instruction in visible text.
The researchers showed that opaque reasoning objects from stronger models could be replayed into weaker compatible models within the same provider family, causing hidden reasoning to be returned in plain text. They documented this across Anthropic Claude, OpenAI GPT, and Google Gemini model families.
The research team analyzed 6,708 public agent transcripts or trajectories from GitHub and Hugging Face and decoded 315,320 hidden reasoning blocks. They recovered sensitive artifacts including PII, API keys, passwords, access tokens, private keys, and other credentials from genuine user sessions.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
3 references tracked. Mallory keeps watching after this page renders.
heise.de
Open sourcethehackernews.com
Open sourcecybersecuritynews.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.