ARTFEED — Contemporary Art Intelligence

GCPO: A New Method to Stabilize Rollout RL for LLMs

ai-technology · 2026-08-13

A new research paper on arXiv (2608.11674) introduces GCPO (Geometrically Constrained Policy Optimization), a method designed to address training instabilities in on-policy rollout reinforcement learning for large language models (LLMs). The paper, titled 'GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs', identifies issues such as cross-task capability degradation and response-length inflation that commonly affect methods like GRPO. The authors propose a dimension-corrected measure called Principal-Subspace Overlap to analyze the stepwise variation of update geometry relative to the dominant singular subspaces of pretrained weights. They observe that transient spikes in overlap often precede performance degradation. GCPO applies hard bilateral orthogonal projections to constrain updates to complementary subspaces, preventing such excursions by construction. The method is evaluated on mathematical reasoning, code generation, and tool-use tasks. The paper is available at https://arxiv.org/abs/2608.11674.

Key facts

  • Paper arXiv:2608.11674 introduces GCPO (Geometrically Constrained Policy Optimization).
  • GCPO addresses training instabilities in on-policy rollout RL for LLMs.
  • Common issues include cross-task capability degradation and response-length inflation.
  • Principal-Subspace Overlap is a new dimension-corrected measure.
  • Transient spikes in overlap often precede performance degradation.
  • GCPO uses hard bilateral orthogonal projections to constrain updates.
  • Evaluated on mathematical reasoning, code generation, and tool-use tasks.
  • Paper available at https://arxiv.org/abs/2608.11674.

Entities

Institutions

  • arXiv

Sources