CoCo: Faithful Response-Level Interpretation for Mixture-of-Experts Reward Models
A recent paper on arXiv (2608.06400) presents Contribution-Contrast (CoCo), a novel approach for interpreting sparse Mixture-of-Experts (MoE) reward models at the response level. The authors contend that routing weights, which indicate which prompts are assigned to an expert, fail to fully explain how experts evaluate responses, offering only a limited view of expert behavior. CoCo defines the roles of experts through selected-rejected response pairs that exhibit the most significant contribution contrasts, effectively capturing both routing and preference behaviors. In both automatic and human assessments, CoCo demonstrates more coherent, faithful, and specialized interpretations compared to router-based, score-based, and sparse autoencoder benchmarks. The research team behind this work aims to enhance the interpretability of reward models essential for aligning AI with human preferences.
Key facts
- Paper on arXiv: 2608.06400
- Introduces Contribution-Contrast (CoCo) method
- Targets sparse Mixture-of-Experts (MoE) reward models
- Uses chosen-rejected response pairs with largest contribution contrasts
- Captures both routing and preference behavior
- Evaluated with automatic and human evaluations
- Outperforms router-based, score-based, and sparse autoencoder baselines
- Aims to improve interpretability of reward models in AI alignment
Entities
Institutions
- arXiv