Rubrics as Privileged Information for Open-Ended Generation
A new arXiv paper (2608.02948) proposes using rubrics as soft privileged information (PI) to improve open-ended generation through on-policy self-distillation (OPSD). The authors argue that while OPSD has been effective in verifiable domains like math, where hard PI (ground-truth answers) constrains valid outputs, open-ended tasks require a different approach. Rubrics, which specify preference structures but admit many valid responses, offer a richer training signal than hard reference completions. The paper demonstrates that distilling towards a single reference completion over-constrains the student model, whereas rubrics capture the shared preference structure across valid responses. This research, submitted to arXiv on August 29, 2026, extends OPSD to domains where multiple correct answers exist, potentially impacting AI-driven creative and generative tasks.
Key facts
- Paper arXiv:2608.02948 proposes rubrics as soft privileged information for open-ended generation.
- On-policy self-distillation (OPSD) uses a single model as both student and teacher.
- OPSD has shown promise in verifiable domains like math with hard PI (ground-truth answers).
- Rubrics guide preferences but admit many valid responses.
- Soft rubric PI provides a larger and more effective training signal than hard reference completion PI.
- Distilling towards a reference completion over-constrains the student.
- Rubrics specify the preference structure shared across valid responses.
- The paper was announced as a cross-type submission on arXiv.
- The research is relevant to open-ended generation tasks.
- The paper is dated August 29, 2026.
Entities
Institutions
- arXiv