SAF-OPD: Stable Advantage Fusion for On-Policy Distillation in RLVR
A recent study introduces SAF, which stands for Stable Advantage Fusion, aimed at merging reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training large language models. The research, titled 'SAF-OPD: Stable Advantage Fusion for On-Policy Distillation,' can be found on arXiv under ID 2607.29209. The authors note that RLVR sends a uniform response-level reward to all tokens, while OPD offers detailed token-level advantages from a superior teacher but limits performance to the teacher's quality and hinders exploration. The naive combination of these methods with a fixed coefficient results in entropy collapse due to mismatches in magnitude and timing. SAF addresses these challenges through a streamlined, four-stage process focused solely on OPD advantages, incorporating a sparsify-then-compress strategy. This paper also suggests a potential conference presentation, tackling a significant issue in aligning language models with verifiable rewards, which could enhance training efficiency and model performance.
Key facts
- Paper title: SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
- arXiv ID: 2607.29209
- Announcement type: cross
- Proposes SAF framework to fuse RLVR and OPD advantages
- Identifies magnitude and temporal mismatches causing entropy collapse
- SAF uses a four-stage pipeline with sparsify-then-compress mechanism
- RLVR broadcasts response-level reward to every token
- OPD scores each token against a stronger teacher for dense advantage
- OPD caps performance at teacher quality and discourages exploration
- Fixed coefficient fusion leads to entropy collapse
Entities
Institutions
- arXiv