Routing-Based On-Policy Distillation Enhances LLM Safety Realignment
A recent study published on arXiv (2607.27081) presents Routing-based On-Policy Distillation (ROPD), a novel framework aimed at enhancing safety realignment in large language models (LLMs) post fine-tuning. While fine-tuning can incorporate detrimental behaviors from harmful data sources, it often retains professional competencies. Current protective measures face issues such as catastrophic forgetting of specialized abilities, failure when unfamiliar prompt templates are used, and susceptibility to re-jailbreaking through changes in system prompts. Instead of adhering to specific templates, ROPD focuses on modeling the differences between aligned and compromised output probability distributions, providing a stronger solution. The paper features comprehensive experiments that evaluate ROPD against alternative strategies.
Key facts
- arXiv paper 2607.27081 proposes Routing-based On-Policy Distillation (ROPD)
- ROPD addresses safety vulnerabilities in LLMs after fine-tuning
- Fine-tuning can embed harmful behaviors from malicious data providers
- Existing defenses cause catastrophic forgetting of specialized skills
- Existing defenses fail when attacker's prompt template is unknown
- Realigned models remain susceptible to re-jailbreaking via system prompt switches
- ROPD models divergence between aligned and compromised output probability distributions
- Extensive experiments compare ROPD against other methods
Entities
Institutions
- arXiv