Scaling Laws for AI Safety: Balancing Character and Rules
A new arXiv paper (2608.13345) introduces a formal model for optimizing AI safety design as deployment scales. The authors propose a resource allocation parameter alpha between character shaping (e.g., RLHF, Constitutional AI) and rule enforcement (e.g., output filters), incorporating scale-dependent filter degradation, common-mode failures, and character fragility. Using a multiplicative Pareto damage model, they derive closed-form expected harm and supplement with tail-risk (CVaR) analysis via Monte Carlo simulation across three scenarios. The paper addresses a gap in formal analysis of how the optimal balance between training-time and inference-time safety measures should change with scale.
Key facts
- Paper arXiv:2608.13345v1, announced as new
- Introduces a stylized comparative-statics model for AI safety design
- Parameterizes safety design as resource allocation alpha in [0,1] between character shaping and rule enforcement
- Incorporates scale-dependent filter degradation, common-mode failures, and character fragility
- Uses multiplicative Pareto damage model to derive closed-form expected harm
- Supplemented with tail-risk (CVaR) analysis via Monte Carlo simulation
- Analyzes three scenarios: optimistic, mod... (incomplete in source)
- Published on arXiv, URL: https://arxiv.org/abs/2608.13345
Entities
Institutions
- arXiv