Anthropic published research showing that reasoning models do not always reveal their full internal deliberations in their visible chain-of-thought, and separately described Claude’s “J-space” as a compact internal workspace tied to deliberate multi-step reasoning. Using a technique referred to as J-lens, the company said it can approximate token-level concepts active inside that workspace before the model produces output or takes actions, offering a view into internal representations that are normally hidden from users and defenders.
Security practitioners said the finding could create a new layer of AI-agent telemetry beyond outputs and action logs, because internal concepts such as “ERROR,” “injection,” “fake,” “fictional,” “manipulation,” and “fraud” reportedly appeared in Claude’s workspace before harmful or deceptive behavior became externally visible. The reporting argues that such visibility could support earlier detection, interruption of multi-step agent workflows, staging-time screening, and post-incident forensics, but notes that this telemetry remains controlled by Anthropic rather than exposed to enterprise customers running agents in production.

Track how attackers are adapting to this technology.
2 events from the most recent confirmed update back to the earliest known activity.
Mitiga published an analysis arguing that Anthropic's J-space and J-lens research could provide a new layer of telemetry for AI agent security by exposing internal model representations before outputs or actions occur. The article says this capability could support detection, interruption of multi-step workflows, staging-time screening, and post-incident forensics, while noting such visibility is not available to enterprise customers.
Anthropic published research titled "Reasoning models don't always say what they think," examining how model internal reasoning can differ from what is expressed externally. The later Mitiga article describes this as newly published interpretability research related to Claude's internal "J-space."
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
2 references tracked. Mallory keeps watching after this page renders.
Map indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.