New Data Valuation Pipeline Improves LLM Alignment
A recent study presents a scalable data valuation framework for Large Language Models (LLMs) that estimates the Shapley value without the need for repeated retraining. This technique tackles data quality challenges in alignment by converting semantic k-NN neighborhoods into a directed graph and assessing data utility via a reference LLM's probability distribution, utilizing zero-shot and one-shot conditional log-likelihood shifts. The framework converts predictive influence scores into localized advantage metrics to identify problematic entries, such as concealed contradictions, safety concerns, and annotation mistakes. Designed for extensive preference and instruction-tuning datasets, this method operates solely on inference.
Key facts
- arXiv paper ID: 2607.22766
- Announce type: cross
- Method approximates Shapley value without iterative retraining
- Uses semantic k-NN neighborhoods mapped into a directed graph
- Evaluates data utility via reference LLM's probability distribution
- Applies zero-shot and one-shot conditional log-likelihood shifts
- Translates influence scores into localized advantage metrics
- Addresses data quality issues in LLM alignment
Entities
Institutions
- arXiv