FAR.AI Introduces Minimal Standard for AI Safeguards with Jailbreak Benchmark
The FAR.AI Minimal Standard for Safeguards, Version 1.0, has been released, providing a taxonomy of 67 static jailbreak techniques and a method for composing them into a large attack space. The standard includes a benchmark of flagship AI models—Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro, and Grok 4.5—evaluated on two datasets totaling 360 attacker goals across CBRNE threats and offensive cyber. A three-stage funnel identifies universal jailbreaks: single prompt templates that elicit compliant responses on over 75% of a domain's goals. The work introduces a cost-to-jailbreak metric to model attacker expenditure. The paper, available on arXiv (2608.03070), aims to provide public evidence on the effectiveness of layered safeguards across developers.
Key facts
- FAR.AI Minimal Standard for Safeguards, Version 1.0 introduced.
- Taxonomy includes 67 static jailbreak techniques.
- Benchmark evaluates Claude Fable 5, GPT-5.6 Sol, Gemini 3.1 Pro, and Grok 4.5.
- Two datasets total 360 attacker goals.
- Goals cover CBRNE threats and offensive cyber.
- Three-stage funnel identifies universal jailbreaks.
- Universal jailbreaks elicit compliant responses on over 75% of domain goals.
- Cost-to-jailbreak metric models attacker spend.
Entities
Institutions
- FAR.AI
- arXiv