ARTFEED — Contemporary Art Intelligence

AI Benchmarks Criticized for Failing to Measure Human-Like Cognition

ai-technology · 2026-08-13

A recent study published on arXiv on February 28, 2025, critiques existing evaluation frameworks for AI, claiming they fall short in measuring human-like cognitive abilities. The authors highlight several significant issues, including the absence of labels validated by humans, insufficient representation of human variability and uncertainty in responses, and a dependence on overly simplistic and ecologically invalid tasks. They performed a human evaluation analysis across nine current AI benchmarks, uncovering substantial flaws in both task and label design. The paper offers five specific recommendations aimed at enhancing future AI evaluation processes, ultimately striving for a deeper and more accurate understanding of AI's resemblance to human cognitive functions. This research is crucial for the AI-technology industry, as it questions the reliability of existing benchmarks and advocates for better assessment methods.

Key facts

  • Paper submitted to arXiv on February 28, 2025.
  • Authors argue current AI evaluation paradigms are insufficient for assessing human-like cognitive capabilities.
  • Key shortcomings include lack of human-validated labels, inadequate representation of human response variability and uncertainty, and reliance on simplified tasks.
  • Human evaluation study conducted on nine existing AI benchmarks.
  • Study suggests major limitations in task and label designs.
  • Five concrete recommendations proposed for future AI evaluation.
  • Paper is in the field of Computer Science and Artificial Intelligence.
  • The paper is titled 'On Benchmarking Human-Like Intelligence in Machines'.

Entities

Institutions

  • arXiv

Sources