RAPO: A Dual-Channel Framework for Risk-Aware Continual Reinforcement Fine-Tuning
A new arXiv preprint (2608.03660) introduces Risk-Aware Policy Optimization (RAPO), a dual-channel framework designed to address catastrophic forgetting in continual reinforcement fine-tuning (RFT) of multimodal large language models. The authors argue that while RFT is generally believed to resist forgetting, it fails under pronounced task distributional shifts, leading to uncontrolled optimization risk. RAPO operates on two channels: the policy channel uses Risk-Aware Policy Scaling to adaptively calibrate per-sample update magnitude based on rollout reliability and Fisher-inspired local predictive sensitivity; the data channel employs Risk-Aware Dynamic Bucket Sampling to reorganize training batches via dynamic risk stratification. The method is presented as a plug-and-play strategy requiring minimal integration, aiming to steer optimization toward informative yet stable samples. The paper is available on arXiv under the identifier 2608.03660.
Key facts
- arXiv paper 2608.03660 proposes RAPO
- RAPO is a dual-channel framework for continual RFT
- Addresses catastrophic forgetting in multimodal large language models
- Policy channel: Risk-Aware Policy Scaling
- Data channel: Risk-Aware Dynamic Bucket Sampling
- Uses rollout reliability and Fisher-inspired local predictive sensitivity
- Plug-and-play strategy
- Published on arXiv
Entities
Institutions
- arXiv