ARTFEED — Contemporary Art Intelligence

Tiny Transformers Show Protoreasoning in Chain of Thought

ai-technology · 2026-08-06

A new paper on arXiv (2608.04980) demonstrates that tiny transformer models, with around 1 million parameters, can effectively use a simplified form of Chain of Thought (CoT) reasoning, termed 'protoreasoning.' This approach allows researchers to study step-by-step reasoning at a much smaller scale than typical large language models (LLMs), enabling more detailed experimentation and analysis. The study addresses the question of whether LLMs learn general algorithms or merely 'bags of heuristics' by using reasoning-friendly tasks based on Dyck languages, which consist of correctly nested brackets. The findings show that protoreasoning traces significantly improve out-of-distribution generalization, and ablations confirm that the content of the trace, not just its existence, is crucial. This work opens up opportunities for more granular investigation into the mechanisms of reasoning in AI systems.

Key facts

  • Paper on arXiv: 2608.04980
  • Tiny transformers with ~1M parameters use protoreasoning
  • Protoreasoning is a simple form of Chain of Thought
  • Study uses Dyck languages (correctly nested brackets)
  • Protoreasoning traces close the out-of-distribution generalization gap
  • Ablations confirm trace content matters
  • Enables detailed experimentation at small scale
  • Addresses generality of reasoning in LLMs

Entities

Institutions

  • arXiv

Sources