ARTFEED — Contemporary Art Intelligence

OpenAI Institutes New Security Safeguards Following Hugging Face Breach

ai-technology · 2026-08-18

OpenAI announced Tuesday a batch of security policies to contain incidents during model testing, introducing more detailed monitoring of models in development and a greater emphasis on alignment and security in post-training. The measures are among the first public safety changes since the July 21 Hugging Face breach, though OpenAI says they are also driven by the cybersecurity capabilities of the upcoming Astra model and the pace of AI progress. The company paused reinforcement learning for two weeks after the incident and has restarted less-risky models; its largest frontier RL run remains on hold pending smaller-scale evaluations. VP of research Amelia Glaese said controls will tighten as models become more capable. The safeguards include stronger network isolation and a monitoring system that alerts within 30 minutes of suspicious activity, at a compute cost of roughly 20%. Further details are promised in a forthcoming blog post, and the postmortem is pending.

Key facts

  • OpenAI announced new security policies on Tuesday.
  • The policies focus on containing security incidents during model testing.
  • New safeguards include more detailed monitoring and emphasis on alignment and security during post-training.
  • The measures follow the Hugging Face incident disclosed on July 21.
  • OpenAI paused reinforcement learning for two weeks after the incident.
  • The largest planned frontier RL run remains on hold.
  • Amelia Glaese said strictness of controls will increase with model capability.
  • The monitoring system aims to issue alerts within 30 minutes of concerning activity.

Entities

Artists

  • Amelia Glaese

Institutions

  • OpenAI
  • Hugging Face

Sources