Researchers disclosed a flaw in encrypted reasoning objects used by major AI reasoning APIs, including OpenAI, Anthropic, and Google, that allowed hidden model reasoning to be reconstructed by replaying opaque reasoning blobs across sessions and, in some cases, across compatible models. The attack did not break encryption directly; instead, it exploited the fact that receiving systems would accept reused encrypted reasoning artifacts, letting weaker or more easily jailbroken models act as approximate decoders for stronger models’ concealed reasoning. Reported risks included extraction of proprietary reasoning, recovery of sensitive data from shared traces, reconstruction of harmful hidden material, and prompt injection hidden inside encrypted reasoning objects.
A practical reproduction showed the technique working against OpenAI models across separate sessions and accounts, recovering semantic reasoning content and even a password present in an original trace. Researchers said they decoded 315,320 reasoning blocks from 6,708 public agent trajectories and recovered hundreds of privacy-related artifacts, including API keys, passwords, access tokens, private keys, and other PII. The issue was disclosed to OpenAI, Anthropic, Google, Microsoft, and Hugging Face, and the primary extraction method was reported to be no longer reproducible after mitigations, although testing indicated inconsistent behavior over time. Researchers said they found no evidence of active exploitation, but warned developers not to publish or retain raw agent traces containing opaque reasoning fields.

Track how attackers are adapting to this technology.
7 events from the most recent confirmed update back to the earliest known activity.
The reproduction author said the attack temporarily stopped working on Tuesday, August 11, at around 6 p.m. PT during testing. This suggested backend or model behavior changes affecting the replay-based recovery method.
The researchers stated that the demonstrated attacks stopped functioning after mitigation efforts, and that the primary extraction technique was no longer reproducible as of August 2026. The report said there was no indication of malicious exploitation in the wild.
According to the later reproduction write-up, Matthew Green showed in May 2026 that encrypted LLM reasoning blobs could be replayed across sessions and accounts, and across models for OpenAI. This established that opaque reasoning artifacts were accepted outside their original context.
The Embrace The Red author implemented the attack against OpenAI by creating a reasoning trace with GPT-5.6 Sol under one account and replaying its encrypted content to GPT-5.6 Luna from another account. The recovered output exposed semantic reasoning content, including a password embedded in the original hidden trace.
The reproduction author reported that the replay-based reasoning recovery attack started working again three days after its August 11 failure. The author described inconsistent reliability while testing the technique against OpenAI models.
In the study "Stealing Reasoning Traces from Proprietary LLM APIs," researchers analyzed 6,708 public agent trajectories, decoded 315,320 reasoning blocks, and recovered hundreds of privacy-related artifacts including API keys, passwords, access tokens, and private keys. The paper also documented proof-of-concept risks such as proprietary reasoning extraction and hidden prompt injection carried inside encrypted reasoning objects.
Researchers said they reported the encrypted reasoning object weakness to OpenAI, Anthropic, Google, Microsoft, and Hugging Face. The disclosure concerned replayable reasoning artifacts that could enable hidden reasoning reconstruction and related risks.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
embracethered.com
Open sourcecysecurity.news
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.