Anthropic's Opus 5 Shows Strong Resistance to Prompt Injection
Anthropic's Opus 5 model demonstrates significantly improved resistance to prompt injection attacks, according to Boris Cherny. In a system card section on page 73, Cherny highlights that across prompt injection evaluations and red teaming, Opus 5 is very difficult to successfully prompt inject. This marks a notable advancement in AI safety, as prompt injection remains a critical vulnerability in large language models. The claim was shared by Simon Willison on 25th July 2026.
Key facts
- Opus 5 is Anthropic's least prompt injectable model yet
- Claim made by Boris Cherny in the system card on page 73
- Resistance confirmed across PI evals and red teaming
- Information shared by Simon Willison on 25th July 2026
Entities
Artists
- Boris Cherny
- Simon Willison
Institutions
- Anthropic