ARTFEED — Contemporary Art Intelligence

DASH: Enhancing On-Policy Self-Distillation for Reasoning Models

other · 2026-08-07

A new arXiv paper (2608.06243) introduces DASH (Divergence-Adaptive Supervision Horizons), a method to improve on-policy self-distillation (OPSD) for training reasoning models. OPSD mitigates the sparsity of verifiable rewards by using a privileged teacher to provide dense token-level supervision at student-visited prefixes. However, standard OPSD assigns equal coefficients to all local divergences, ignoring the temporal context of the rollout. DASH adapts supervision horizons based on divergence sequences, allowing the model to better exploit the temporal structure of autoregressive generation. The paper is authored by researchers and was announced on arXiv. The method aims to enhance the reasoning capabilities of large language models by more effectively utilizing dense supervision signals.

Key facts

  • Paper ID: arXiv:2608.06243
  • Announce Type: new
  • Method: DASH (Divergence-Adaptive Supervision Horizons)
  • Improves on-policy self-distillation (OPSD)
  • Addresses sparsity of verifiable rewards in RLVR
  • Uses privileged teacher for dense token-level supervision
  • Standard OPSD assigns equal coefficients to local divergences
  • DASH adapts supervision horizons based on divergence sequences

Entities

Institutions

  • arXiv

Sources