ARTFEED — Contemporary Art Intelligence

CogTest Benchmark Evaluates Cognitive Habits in Large Reasoning Models

ai-technology · 2026-07-29

A new evaluation tool named CogTest has been launched to assess cognitive habits in Large Reasoning Models (LRMs). This benchmark draws inspiration from the way LRMs reflect human-like cognitive patterns, such as the self-checking phrase, 'Wait, did I miss anything?' CogTest is rooted in the Habits of Mind framework, which outlines cognitive habits linked to effective human problem-solving. It comprises 16 cognitive habits, each represented by 25 varied tasks, employing an evidence-first extraction approach for accurate habit detection. The benchmark assessed 16 prominent LLMs, consisting of 13 LRMs and 3 non-reasoning models. Results indicate that LRMs exhibit cognitive habits akin to those of humans, unlike traditional LLMs. The research is documented in the paper 'Towards Understanding the Cognitive Habits of Large Reasoning Models,' accessible on arXiv with identifier 2506.21571.

Key facts

  • CogTest is a benchmark for evaluating cognitive habits in LRMs.
  • It includes 16 cognitive habits, each with 25 tasks.
  • The benchmark uses an evidence-first extraction method.
  • 16 LLMs were evaluated: 13 LRMs and 3 non-reasoning models.
  • LRMs exhibit human-like cognitive habits, unlike conventional LLMs.
  • The paper is on arXiv with identifier 2506.21571.
  • Cognitive habits are based on the Habits of Mind framework.
  • Phrases like 'Wait, did I miss anything?' are observed in LRMs.

Entities

Institutions

  • arXiv

Sources