ARTFEED — Contemporary Art Intelligence

TCPO: Turn-Level Credit Policy Optimization for Verifier-Guided Multi-Turn RL

ai-technology · 2026-08-04

A new arXiv paper (2608.01667) introduces TCPO, a turn-level credit assignment method for verifier-guided multi-turn reinforcement learning in large language models. The method addresses the gap between dense feedback and dense credit by converting verifier scores into turn-level advantages through reference-based comparisons. TCPO includes retrospective credit for immediate progress, hindsight delayed credit for non-improving turns with later payoff, and selective fixed-history counterfactual estimation. The paper is authored by researchers and was announced on arXiv.

Key facts

  • Paper arXiv:2608.01667 proposes TCPO, a turn-level credit assignment method.
  • TCPO is designed for verifier-guided multi-turn reinforcement learning.
  • It converts verifier scores into turn-level advantages.
  • Retrospective credit captures immediate progress and regression.
  • Hindsight delayed credit identifies non-improving turns with later payoff.
  • Selective fixed-history counterfactual estimation refines high-surprisal turns.
  • The paper is available at https://arxiv.org/abs/2608.01667.
  • The announcement type is 'new'.

Entities

Institutions

  • arXiv

Sources