MIST and SCOPE: New Framework for Selective Trust in Language Models
A recent study published on arXiv (2608.06377) presents MIST, a benchmark created through human annotation to assess how language models selectively trust external signals. This benchmark evaluates reasoning items across four distinct scenarios: clean, misleading, correct-context, and irrelevant-context. The researchers introduce SC2W, a metric that tracks the frequency at which a misleading signal alters a clean-correct answer to incorrect. Their extensive benchmark analysis reveals that vulnerability to misleading signals is a widespread issue. To combat this, they introduce SCOPE, which identifies clean-correct/misleading-wrong failures and enhances a standard Direct Preference Optimization (DPO) objective using balanced preference pairs across all four scenarios. This paper is noted as a cross submission and can be accessed via the provided URL.
Key facts
- Paper on arXiv:2608.06377
- Introduces MIST, a human-annotated benchmark
- Four matched conditions: clean, misleading, correct-context, irrelevant-context
- SC2W metric counts misleading flips
- Susceptibility to misleading signals is universal
- SCOPE uses DPO objective
- Balanced preference pairs across conditions
- Announce type: cross
Entities
Institutions
- arXiv