CogArena Benchmark Evaluates LLM Cognitive Ability Structure
Researchers have unveiled CogArena, a benchmark featuring 13 paradigms designed to assess cognitive ability structures in large language models (LLMs) through procedural generation. This benchmark employs a multimethod approach to identify when cognitive-task scores should be assigned dimensional labels within five theory-driven categories. An analysis involving 55 open-weight models revealed that nearly all correlations among paradigms were positive, with a shared axis accounting for approximately half of the variance. The advantage within groupings was minimal, sensitive to scoring, and varied among model families. Additionally, a distinct study involving 12 models from six families indicated a slight matched-grouping benefit with specific scaffolds, but no scaffold-specific differences persisted after multiplicity correction, questioning the notion of neatly categorizing LLM cognitive abilities.
Key facts
- CogArena is a procedurally generated 13-paradigm benchmark.
- The benchmark evaluates cognitive ability structure in LLMs.
- 55 open-weight models were tested.
- Nearly all paradigm correlations were positive.
- A common axis explained about half the variance.
- The within-grouping advantage was small and uncertain.
- A frozen study involved 12 models from six families.
- No scaffold-specific contrast survived multiplicity correction.
Entities
Institutions
- arXiv