Anthropic and OpenAI Propose Embedding Independent Safety Evaluators, But Details Remain Unclear
Dario Amodei, the CEO of Anthropic, has suggested that third-party evaluators should be integrated into leading AI firms to monitor safety incidents and evaluate model alignment. Anthropic will allow organizations such as METR and Redwood Research to access its systems, with support from OpenAI's CEO, Sam Altman. These evaluators aim to obtain access to intermediate training checkpoints and conduct employee interviews. Adam Gleave from FAR.AI mentioned this could uncover troubling behaviors, while Alexander Meinke of Apollo Research highlighted the importance of verifying AI alignment. Details about the evaluators and their access have not been revealed by Anthropic or OpenAI. California's SB 53 and SB 813 require safety frameworks and independent checks, while the EU AI Act mandates model evaluations. Amodei's proposal goes beyond existing regulations.
Key facts
- Anthropic CEO Dario Amodei proposed embedding third-party evaluators inside all frontier AI companies in an essay published over the weekend.
- Anthropic committed to giving independent evaluators like METR and Redwood Research unprecedented access to its systems.
- OpenAI CEO Sam Altman said OpenAI would also commit to the practice.
- Evaluators want access to intermediate training checkpoints, post-training environments, evaluation transcripts, and employee interviews.
- Neither Anthropic nor OpenAI has shared which evaluators they will work with, when they will be embedded, how many, what systems they can access, or what can be disclosed publicly.
- During the Hugging Face incident, OpenAI gave METR and Redwood roughly a week on premises to investigate.
- For GPT-6 Astra, Apollo Research was given only three days to test the model.
- California's SB 53 requires large frontier AI developers to publish safety frameworks and report critical safety incidents.
- California's SB 813 creates a framework for state-recognized independent verification organizations.
- The EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents.
Entities
Artists
- Dario Amodei
- Sam Altman
- Alexander Meinke
- Adam Gleave
- John Steidley
- Henry Papadatos
- Demis Hassabis
- Jacob Coxon
- Brian Merchant
- Elon Musk
- Evan Hubinger
- Donald Trump
- Kevin Hassett
- Chris Lehane
- David Sacks
- Daniela Amodei
Institutions
- Anthropic
- OpenAI
- METR
- Redwood Research
- TechCrunch
- Apollo Research
- FAR.AI
- Palisade Research
- Safer AI
- Meta
- SpaceXAI
- Google DeepMind
- EU AI Office
- Volkswagen
- HuggingFace
- US government
- xAI
- All-In Summit
- CNBC
- Truth Social
- X
- National Economic Council
- Chinese Foreign Ministry
- White House
- Bloomberg
- Fortune
- The Information
- U.S. government
- FRONTIER Act
- Google-DeepMind
- Freitag
Locations
- California
- United States
- Europe
- China
- Germany
- Los Angeles
- Washington
- Vereinigte Staaten
Sources
- TechCrunch AI —
- Quartz —
- TechCrunch AI —
- der Freitag —
- SCMP Culture —
- TechCrunch AI —