Anthropic Makes Auto Mode Default in Claude Code, Citing Safety Evals
Starting August 14th, Anthropic will implement auto mode as the standard setting for new sessions in Claude Code for Pro, Max, and Team plans. This change stems from the company's strong belief in the mode's safety, as highlighted during a Fireside Chat with Cat Wu and Thariq Shihipar at last month's AI Engineer World’s Fair. Wu noted that nearly all Anthropic employees utilize auto mode, emphasizing that they have addressed significant attack vectors like prompt injection and data exfiltration, resulting in lower risks compared to typical human reviewers. Evaluations showed that in a test with 1,053 paid participants, only 13.6% rejected a harmful command, whereas auto mode would have blocked 89%. A Trajectory Labs assessment found no success in 720 indirect prompt injection attempts against Claude Fable 5, Opus 5, and Sonnet 5. Nonetheless, concerns persist regarding attacks from malicious packages that could instruct agents to execute harmful commands, which auto mode might not stop. Simon Willison, the post's author, urges for further independent validation while remaining cautiously optimistic.
Key facts
- Auto mode becomes default in Claude Code for Pro, Max, and Team plans on August 14th.
- Cat Wu and Thariq Shihipar discussed auto mode at the AI Engineer World’s Fair.
- Anthropic claims to have mitigated prompt injection and data exfiltration risks.
- In a test with 1,053 paid testers, only 13.6% of humans refused a dangerous command, while auto mode would have blocked 89%.
- Trajectory Labs evaluated Claude Fable 5, Opus 5, and Sonnet 5; none of 720 indirect prompt injection attacks succeeded.
- Simon Willison highlights potential vulnerabilities like malicious packages that auto mode may not catch.
- Willison predicts a 'challenger disaster for coding agents security' in 2026.
- The post was published on August 8, 2026.
Entities
Institutions
- Anthropic
- Trajectory Labs
- AI Engineer World’s Fair