New Framework for Evaluating Text Evaluation Metrics
A new arXiv paper (2608.01423) proposes a framework for assessing reference-based text evaluation metrics, which are used to score candidate responses in natural language generation by comparing them to reference responses. The authors argue that while statistical correlation with human ratings is the standard measure of reliability, it is insufficient when metrics are used as optimization objectives, as agents may strategically game the metrics. They introduce two complementary notions of alignment: statistical alignment (correlation with human ratings) and strategic alignment (resistance to perturbations that do not add task-relevant information). The paper contributes test principles for reference-based metrics, including human-rating correlation, degradation sensitivity, and manipulation robustness, to evaluate whether a metric agrees with human judgments and resists gaming. The paper is announced as a new submission on arXiv.
Key facts
- Paper arXiv:2608.01423 proposes a framework for evaluating text evaluation metrics.
- Introduces statistical alignment and strategic alignment.
- Proposes test principles: human-rating correlation, degradation sensitivity, and manipulation robustness.
- Argues correlation alone is insufficient for metrics used as optimization objectives.
- Focuses on reference-based metrics in natural language generation.
- Addresses the issue of agents strategically gaming evaluation metrics.
- Published on arXiv with announcement type 'new'.
Entities
Institutions
- arXiv