A proof-of-concept indirect prompt-injection attack reportedly induced Claude Code Opus 5 in Auto Mode to execute attacker-controlled code after processing a malicious website during an ostensibly benign summarization task. The chain caused the agent to fall back from WebFetch to curl, download and extract a ZIP archive, then create and run a Python decoder inside the attacker-controlled extraction directory. A bundled struct.py abused Python import-path precedence to shadow the standard-library struct module when the decoder imported Base64-related functionality, enabling a staged payload to contact attacker-controlled command-and-control infrastructure and launch Calculator.
The researcher reported success rates of 60% for Python/C2 and nested Claude subprocess variants, and 80% for a variant that opened Calculator and wrote outside the designated workspace, in limited testing. Anthropic reportedly categorized the disclosure as Informative, stating that Auto Mode is a convenience capability governed by a best-effort classifier rather than a security boundary. Organizations running autonomous coding agents against untrusted content should isolate the runtime, restrict outbound network access, withhold sensitive credentials, constrain filesystem access, and monitor agent-created processes and tool calls.

Track how attackers are adapting to this technology.
5 events from the most recent confirmed update back to the earliest known activity.
According to researcher Johann Rehberger, Claude Code began using Auto Mode by default in mid-August 2026. Auto Mode uses an automated classifier in place of user confirmations for potentially dangerous actions.
The researcher submitted the issue through Anthropic's model bug-bounty email and security reporting channel. Anthropic reportedly closed it as Informative, stating that Auto Mode is a convenience feature with a best-effort classifier rather than a security boundary, and that OS isolation and network-egress controls are the relevant protections.
In a variant of the proof of concept, the malicious import launched a headless nested Claude Code instance via claude -p. The nested agent reportedly ran whoami, uname, and id, opened Calculator, and wrote files under the user's home directory or outside the original workspace.
Embrace The Red demonstrated a lab proof of concept in which an attacker-controlled website induced Claude Code Opus 5 Auto Mode to use curl, download a redirected ZIP archive, and execute a self-written Python decoder from the extracted directory. A malicious struct.py shadowed Python's standard-library module, enabling a staged payload to call back to a controlled C2 server and open Calculator; reported success rates ranged from 60% to 80% across variants.
A third-party evaluation commissioned by Anthropic and conducted by Trajectory Labs reportedly recorded a 0.00% prompt-injection attack success rate for Claude Code Opus 5 Auto Mode across 72 scenarios, each run 10 times.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
5 references tracked. Mallory keeps watching after this page renders.
xakep.ru
Open sourcetheregister.com
Open sourcecybersecuritynews.com
Open sourcecyberveille.ch
Open sourceembracethered.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.