ARTFEED — Contemporary Art Intelligence

ADAGE: New Benchmark Exposes Cultural Reasoning Gap in AI

ai-technology · 2026-07-29

A team of researchers has introduced ADAGE (Analogical Difficulty-by-design Assessment for Grounded Evaluation), a versatile framework designed to establish translation-free benchmarks aimed at assessing abstract analogical reasoning in artificial intelligence. This framework merges curation by native speakers with generation assistance from large language models to produce benchmarks that are culturally relevant. Benchmarks have been developed for Arabic, Amharic, and Japanese. An assessment of 14 open-weight models uncovered a notable cultural reasoning disparity: models that excelled in English proverb reasoning experienced accuracy declines ranging from 12 to 52 percentage points on the native benchmarks. The framework, along with the benchmarks and evaluation suite, is now available to the public.

Key facts

  • ADAGE is a language-agnostic pipeline for analogical reasoning evaluation.
  • It combines native-speaker curation with LLM-assisted generation.
  • Benchmarks were constructed for Arabic, Amharic, and Japanese.
  • 14 open-weight models were evaluated.
  • Accuracy dropped by 12–52 percentage points on native benchmarks compared to English.
  • The pipeline and benchmarks are publicly released.
  • The research addresses the problem of multilingual reasoning evaluation relying on translated English benchmarks.
  • The study highlights a cultural reasoning gap in AI models.

Entities

Institutions

  • arXiv

Sources