Anthropic said Claude Code Auto Mode will become the default for Pro, Max, and Team users after testing showed people routinely approve risky coding-agent actions. Internal results cited by the company found users approve 97% of permission prompts, while a study with more than 1,000 participants found humans detected only 13.6% of dangerous commands overall and just 5% after 50 prompts; Anthropic said Auto Mode detected about 89% of those risky actions. The company said it has added hard denies for data exfiltration, protections for secrets and confidential data, checks on Git destinations and destructive Git operations, and API-side screening for prompt injection attempts.
Separate testing reinforced the same pattern: a browser-based simulation built by Belgian developer Alex Wauters recorded more than 40,000 game runs and 409,000 approval decisions, with players allowing roughly one in three malicious requests. Scope violations, typosquatted packages, and suspicious curl requests were among the prompts that frequently slipped through, especially under time pressure, highlighting approval fatigue as a core weakness in human review. Anthropic and outside researchers both said automated controls are not a complete fix, but argued that stronger permission models, sandboxing, cloud devcontainers, contextual hooks, and layered defenses are more reliable than depending on repeated human approvals alone.

Track how attackers are adapting to this technology.
9 events from the most recent confirmed update back to the earliest known activity.
Alex Wauters launched a browser-based game in late May to test whether humans can safely approve or deny AI coding-agent permission requests, and published a concurrent blog post about it.
Anthropic launched Claude Code Auto Mode in March 2026 as a research preview before later broader rollout. The feature routes tool calls through a classifier to reduce reliance on constant manual approval prompts.
Anthropic introduced Auto Mode earlier in 2026 to let Claude Code decide when human intervention is needed using a separate classifier model.
Anthropic announced that Claude Code Auto Mode will become the default on August 14 for Pro, Max, and Team users, while remaining opt-in for Enterprise, API, and cloud-platform users for now.
Anthropic said third-party testing by Trajectory Labs found none of 720 prompt-injection attacks succeeded against Anthropic models running in Auto Mode.
In a two-week red-teaming pilot with Apollo Research, Anthropic said the classifier missed 12% of attacks before hardening and 7% afterward on a held-out test set.
Ahead of making Auto Mode the default, Anthropic added hard denies for data exfiltration, rules for secrets and confidential data, git destination and destructive-action checks, and an API-side prompt-injection screening probe.
Anthropic tested more than 1,000 participants and found humans caught 13.6% of dangerous commands overall, dropping to 5% after 50 prompts, while Claude Code in Auto Mode caught 89%.
After more than 40,000 game runs and about 409,000 approval decisions, Wauters found players approved roughly one in three malicious requests, with scope violations, typosquatted packages, and suspicious curl requests often missed.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
Follow how adversaries are adapting to this technology, and where it touches your stack today.
6 references tracked. Mallory keeps watching after this page renders.
infoworld.com
Open sourcetheregister.com
Open sourcehelpnetsecurity.com
Open sourcethenewstack.io
Open sourcetheregister.com
Open sourceclaude.com
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.