ARTFEED — Contemporary Art Intelligence

Relay-OPD: A New Method to Fix Prefix Failure in On-Policy Distillation

ai-technology · 2026-07-29

A new paper on arXiv (2607.26057) introduces Relay On-Policy Distillation (Relay-OPD), a method to address prefix failure in on-policy distillation (OPD). In OPD, a student model learns from a teacher model using its own generation trajectory, but if the student makes an early mistake, subsequent tokens build on that error, leading to unreliable supervision. The authors identify a teacher-student continuation asymmetry on failed prefixes: the teacher tends to redirect, while the student continues along the wrong path. Relay-OPD uses this asymmetry as a label-free trigger for the teacher to briefly take over at critical early positions, producing a teacher leg, after which the student resumes. A limited relay budget concentrates intervention on key points while minimizing policy deviation. The method is evaluated on mathematical reasoning tasks, showing improved performance over standard OPD.

Key facts

  • Paper arXiv:2607.26057 introduces Relay-OPD
  • On-policy distillation suffers from prefix failure
  • Teacher-student continuation asymmetry on failed prefixes is identified
  • Relay-OPD uses a label-free handoff trigger
  • Teacher briefly takes over at detected trigger points
  • Limited relay budget concentrates intervention on early positions
  • Method evaluated on mathematical reasoning tasks
  • Relay-OPD improves over standard OPD

Entities

Institutions

  • arXiv

Sources