OM-GRPO: A New Label-Free RLVR Framework to Prevent Answer Collapse
A recent study published on arXiv (2608.03119) presents OM-GRPO, a framework designed for label-free Reinforcement Learning with Verifiable Rewards (RLVR) that tackles the challenge of answer collapse. Conventional RLVR depends on accurate answers, which hampers scalability. While voting-based label-free approaches substitute gold-standard supervision with consensus from model samples, they risk reinforcing answer tokens instead of enhancing reasoning. OM-GRPO separates reward estimation from policy optimization by masking gradients on the answer span while still using answer-level rewards through a soft consensus signal. Additionally, it introduces Contrast-Augmented Reward, which improves reward estimation via economical pairwise comparisons of existing trajectories without requiring extra rollouts. The paper indicates enhancements across various reasoning benchmarks, although specific results are not provided in the abstract.
Key facts
- Paper arXiv:2608.03119 proposes OM-GRPO, a label-free RLVR framework.
- OM-GRPO masks gradients on the answer span to prevent answer collapse.
- It retains answer-level rewards through a soft consensus signal.
- Contrast-Augmented Reward refines reward estimation via pairwise comparisons.
- The method avoids additional rollouts for reward refinement.
- It addresses limitations of voting-based label-free RLVR methods.
- The paper is announced on arXiv with type 'new'.
- The framework is evaluated on diverse reasoning benchmarks.
Entities
Institutions
- arXiv