Cloud-ScPO: Topology-Guided Preference Mining for LLM Reasoning
A recent study published on arXiv (2608.01014) presents Cloud-ScPO, a framework designed for semi-supervised preference optimization aimed at enhancing mathematical reasoning in large language models (LLMs). The researchers note that reasoning paths for various math problems create organized global point clouds within the model's internal representation space, where correct and incorrect paths are geometrically distinct. Cloud-ScPO employs a limited labeled dataset to generate multiple reference clouds for correct and incorrect trajectories, scoring each path using a soft k-nearest-neighbor approach based on mean-pooled hidden states. This method effectively extracts preference pairs without needing verified answers, human input, or external rewards, addressing a significant challenge in preference optimization. The paper, marked as a replace-cross announcement, adds to the expanding domain of AI interpretability and alignment, showcasing a unique approach that utilizes hidden-state geometry for semi-supervised learning. The name Cloud-ScPO signifies its cloud-based reference construction and scoring system, making it pertinent for AI researchers and practitioners engaged in reasoning and preference learning tasks.
Key facts
- Paper arXiv:2608.01014 introduces Cloud-ScPO.
- Cloud-ScPO is a semi-supervised preference optimization framework.
- It targets mathematical reasoning in large language models.
- The method uses hidden-state geometry to derive preference supervision.
- Correct and incorrect reasoning trajectories show different geometric organization.
- A small labeled set is used to construct reference clouds.
- Trajectories are scored using a component-level soft k-nearest-neighbor measure.
- The approach avoids verified answers, human annotations, and external reward models.
Entities
Institutions
- arXiv