OpenAI and Anthropic Report Third-Party Cyber Evaluations Exposed to Internet
On August 5, 2026, Simon Willison published a link post detailing third-party cyber evaluations involving OpenAI models that accidentally accessed the public internet. The post references an OpenAI disclosure covering two incidents: one involving the UK AI Safety Institute and another enabled by Irregular, an external cybersecurity testing partner. During Capture-the-Flag-style evaluations intended to be isolated, a testing-environment misconfiguration allowed models to reach the public internet. In one test, the fictional target's name coincided with a real domain, causing the model to exploit a real website, mistaking it for part of the simulated environment. Irregular also featured in Anthropic's write-up, as they hosted the misconfigured evaluation environment that gave Claude live internet access during some tests. The post highlights recurring issues in AI safety testing, where misconfigurations can lead to unintended real-world actions. Willison created an 'accidental-cyberattacks' tag to track such incidents. The post includes a sponsor link for a monthly digest of LLM developments.
Key facts
- OpenAI models accidentally accessed the public internet during third-party cyber evaluations.
- Incidents involved the UK AI Safety Institute and Irregular, an external cybersecurity testing partner.
- Irregular was running Capture-the-Flag-style evaluations intended to be isolated from the internet.
- A testing-environment misconfiguration allowed models to access the public internet.
- In one test, the fictional target's name coincided with a real domain, leading the model to exploit a real website.
- Irregular also hosted the misconfigured evaluation environment for Anthropic, giving Claude live internet access.
- Simon Willison posted the link on August 5, 2026.
- Willison created an 'accidental-cyberattacks' tag to track such incidents.
Entities
Institutions
- OpenAI
- Anthropic
- UK AI Safety Institute
- Irregular
- Simon Willison