ARTFEED — Contemporary Art Intelligence

PS-OPSD: New Method Improves Reasoning in Self-Distilled LLMs

ai-technology · 2026-08-04

A recent study published on arXiv (2608.01589) presents Problem-Space-Guided On-Policy Self-Distillation (PS-OPSD), a novel technique aimed at improving reasoning within large language models. Instead of relying on complete reference solutions, this approach utilizes trajectory-based guidance that outlines the initial state, goal conditions, constraints, and a chosen state-transition path. This method addresses the limitations of teacher-provided token-level targets, which often depend on reference-specific data that is not accessible during inference. The student rollout and OPSD objective remain unchanged. In tests across three mathematical reasoning benchmarks and model sizes ranging from 1.7B to 8B, PS-OPSD outperformed other methods in aggregate question-only accuracy. Controlled experiments reveal that the relevance of guidance and coherence of paths significantly enhance performance, emphasizing the importance of representing privileged information. The research team has made this work available on arXiv, contributing to advancements in artificial intelligence, particularly in enhancing language models' reasoning abilities through self-distillation methods.

Key facts

  • PS-OPSD replaces complete solutions with trajectory-grounded guidance.
  • Guidance includes initial state, goal conditions, constraints, and state-transition path.
  • Method tested on three mathematical reasoning benchmarks.
  • Model scales range from 1.7B to 8B parameters.
  • PS-OPSD achieves highest aggregate question-only accuracy.
  • Guidance relevance and path coherence contribute to gains.
  • Paper is on arXiv with ID 2608.01589.
  • Announcement type: new.

Entities

Institutions

  • arXiv

Sources