ARTFEED — Contemporary Art Intelligence

Self-Guided Process Reward Optimization for LLM Reasoning

ai-technology · 2026-07-27

A new framework called Self-Guided Process Reward Optimization (SPRO) has been proposed to enhance the reasoning capabilities of Large Language Models (LLMs) through process reinforcement learning. SPRO eliminates the need for additional process reward models by deriving process rewards intrinsically from the policy model itself. It introduces two key innovations: Cumulative Process Rewards (CPR) and Masked Step Advantage (MSA), which enable rigorous step-wise action advantage estimation within shared-prompt sampling groups. The approach addresses the lack of a unified theoretical framework for process-level advantage estimation and reduces computational overhead. Experimental results demonstrate its effectiveness in improving LLM reasoning.

Key facts

  • SPRO is a framework for process reinforcement learning in LLMs.
  • It derives process rewards from the policy model itself, avoiding extra reward models.
  • It introduces Cumulative Process Rewards (CPR) and Masked Step Advantage (MSA).
  • CPR and MSA enable step-wise action advantage estimation within shared-prompt sampling groups.
  • The framework reduces computational overhead compared to using separate process reward models.
  • It provides a unified theoretical framework for process-level advantage estimation.
  • Experimental results show SPRO enhances LLM reasoning capabilities.
  • The paper is available on arXiv with ID 2507.01551.

Entities

Institutions

  • arXiv

Sources