Claude Code Auto Mode Goes Default Aug 14, Outperforms Human Review on Safety
Anthropic will make auto mode the default permission setting for Claude Code on Pro, Max, and Team plans starting August 14, 2025. Instead of prompting users to approve each file edit or shell command, a background classifier model will silently review actions and only interrupt when it detects something risky, such as scope escalation, credential access, or data exfiltration. In a controlled study with 1,053 paid testers, the classifier caught 89% of injected dangerous commands compared to just 13.6% caught by human reviewers. Anthropic's own usage data shows users approve 97% of routine prompts, with the likelihood of catching a genuinely dangerous command dropping from roughly 17% early in a session to about 5% after 50 or more approvals. Despite outperforming distracted human review, Anthropic acknowledges the system is not risk-free, noting that red-team testing found the classifier still missed 7% of synthetic attacks after hardening.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in