ARTFEED — Contemporary Art Intelligence

arXiv Study Reveals Authenticity Gap in Human Evaluation of Natural Language Generation

ai-technology · 2026-08-19

Human assessments are regarded as the benchmark for assessing natural language generation (NLG) systems; however, an arXiv preprint (ID 2205.11930v3) suggests that this approach frequently misrepresents genuine human preferences. The authors apply utility theory to scrutinize the evaluation protocol, uncovering that underlying assumptions about annotators may distort ratings. They indicate that Likert scales could potentially invert authentic preferences. While they suggest enhancements to the protocol, they acknowledge its limitations in evaluating open-ended tasks, such as story generation. A novel human evaluation methodology is proposed, underscoring the necessity for reliable assessment strategies across diverse tasks. This study raises concerns about the dependability of existing methods and emphasizes the gap between numerical averages and true human evaluations.

Key facts

  • Human ratings are considered the gold standard in NLG evaluation.
  • The standard protocol averages annotator ratings to rank NLG systems.
  • The study applies utility theory from economics to analyze evaluation.
  • Implicit assumptions about annotators are often violated in practice.
  • Likert scales can provably reverse the direction of true preferences in certain cases.
  • The authors propose improvements to the standard protocol.
  • Even the improved protocol cannot evaluate open-ended tasks like story generation.
  • A new human evaluation method is proposed for open-ended tasks.

Entities

Sources