TCPO: Turn-Level Credit Policy Optimization for Verifier-Guided Multi-Turn RL
A new arXiv paper (2608.01667) introduces TCPO, a turn-level credit assignment method for verifier-guided multi-turn reinforcement learning in large language models. The method addresses the gap between dense feedback and dense credit by converting verifier scores into turn-level advantages through reference-based comparisons. TCPO includes retrospective credit for immediate progress, hindsight delayed credit for non-improving turns with later payoff, and selective fixed-history counterfactual estimation. The paper is authored by researchers and was announced on arXiv.
Key facts
- Paper arXiv:2608.01667 proposes TCPO, a turn-level credit assignment method.
- TCPO is designed for verifier-guided multi-turn reinforcement learning.
- It converts verifier scores into turn-level advantages.
- Retrospective credit captures immediate progress and regression.
- Hindsight delayed credit identifies non-improving turns with later payoff.
- Selective fixed-history counterfactual estimation refines high-surprisal turns.
- The paper is available at https://arxiv.org/abs/2608.01667.
- The announcement type is 'new'.
Entities
Institutions
- arXiv