ARTFEED — Contemporary Art Intelligence

ICSD: New Method for On-Policy Self-Distillation in Agentic RL

ai-technology · 2026-08-18

A recent study presents Influence Calibration for Self-Distillation (ICSD), an innovative technique aimed at enhancing on-policy self-distillation (OPSD) within agentic reinforcement learning. This research, which can be found on arXiv (2608.14945), tackles the issue of trust-utility misalignment prevalent in current OPSD approaches that assign token-level supervision based on teacher trust, often neglecting whether the focus on a token aligns with the policy goals. ICSD evaluates the first-order response of the importance-weighted RL surrogate to teacher-induced output changes for each token being supervised. By employing batch-adaptive calibration, it transforms this fluctuating signal into a controlled allocation weight, maintaining the original auxiliary-loss mass during each action turn. The detached weights influence only the distillation loss, eliminating the need for an extra model pass. Evaluations on ALFWorld, WebShop, and Search-QA benchmarks demonstrated improvements across all relevant aggregate metrics compared to trust-only allocation. The study was recently submitted to arXiv.

Key facts

  • ICSD is a new method for on-policy self-distillation in agentic RL.
  • It addresses the trust-utility mismatch in existing OPSD methods.
  • ICSD measures first-order response of RL surrogate to teacher-directed perturbation.
  • Batch-adaptive calibration produces bounded allocation weights.
  • Weights are detached and affect only distillation loss.
  • No additional model pass is required.
  • Evaluated on ALFWorld, WebShop, and Search-QA.
  • Improves all matched aggregate metrics over trust-only allocation.

Entities

Institutions

  • arXiv

Sources