Cat-DPO: Adaptive Safety Alignment for LLMs via Per-Category Margins
A recent paper published on arXiv (2604.17299) presents Cat-DPO, an innovative direct-preference-optimization algorithm that approaches safety alignment as a constrained optimization issue for each category. In contrast to traditional techniques that utilize a uniform safety margin for all preference pairs, Cat-DPO differentiates by applying an adaptive safety margin tailored to each harm category. This margin becomes stricter when unsafe responses persist within a category and loosens as the model demonstrates improvement, enabling the training signal to reflect the unique challenges of each category rather than relying on a single global metric. The authors contend that conventional preference-based safety alignment oversimplifies safety into a single scalar, leading to models that may seem safe overall but are still vulnerable in less common harm categories. Cat-DPO seeks to remedy this by adjusting safety constraints dynamically for each category. The paper includes experiments across two LLM backbones and six preference-learning benchmarks, indicating that Cat-DPO enhances overall helpfulness while ensuring safety, although specific numerical outcomes are not detailed in the abstract. This research is significant for AI safety, especially in refining large language models to effectively balance helpfulness with the refusal of harmful requests. The paper can be accessed on arXiv with the identifier 2604.17299, categorized as 'replace-cross'.
Key facts
- Paper arXiv:2604.17299 introduces Cat-DPO
- Cat-DPO is a direct-preference-optimization algorithm with per-category adaptive safety margins
- It casts safety alignment as a per-category constrained optimization problem
- The margin tightens when the model produces unsafe responses in a category and relaxes when it improves
- Experiments conducted across two LLM backbones and six preference-learning baselines
- Cat-DPO improves aggregate helpfulness while maintaining safety
- The paper addresses the issue of uniform safety margins leading to unsafe minority categories
- The paper is announced as 'replace-cross' on arXiv
Entities
Institutions
- arXiv