Small Foundation Models Match 70B Baselines on Human Behavior Data
A recent study published on arXiv (2608.05224) explores the scaling characteristics of small foundation models developed using human behavioral data. Researchers created fourteen models, varying from 135M to 14B parameters, utilizing four different architectural families on the Psych-101 dataset, which includes 10.7 million trial-level choices from 160 experiments. In terms of in-distribution performance, the results indicated that scale had minimal impact, as models clustered closely together, with those having 0.6B to 1B parameters achieving results comparable to a 70B baseline for held-out participants. Conversely, out-of-distribution performance revealed a steeper scaling gradient, favoring larger models for generalizing to new task structures. To analyze the information utilized by these models, two diagnostics were employed, systematically removing four prompt channels: task instructions, experimental stimuli, outcome feedback, and choice. The results imply that while small models can act as effective cognitive proxies, their ability to generalize is scale-dependent when encountering unfamiliar tasks. This study prompts further inquiry into whether these models understand task structure or rely on statistical shortcuts. The full paper can be accessed at https://arxiv.org/abs/2608.05224.
Key facts
- Study on arXiv:2608.05224
- Fourteen models trained from 135M to 14B parameters
- Four architecture families
- Psych-101 dataset with 10.7 million trial-level choices from 160 experiments
- In-distribution, scale barely matters; models fall within a narrow band
- 0.6B to 1B parameters match a 70B baseline on held-out participants
- Out-of-distribution, larger models show steeper scaling gradient
- Two diagnostics strip four prompt channels: task instructions, experimental stimuli, outcome feedback, and choice
Entities
Institutions
- arXiv