ARTFEED — Contemporary Art Intelligence

Unified Benchmark Reveals Limits of Neural Language Models in Poverty of Stimulus Tests

ai-technology · 2026-08-17

A new study introduces PoSHBench, a unified benchmark for evaluating the Poverty of the Stimulus Hypothesis (PoSH) in artificial neural networks (ANNs). The research, published on arXiv, tests Transformer, LSTM, and n-gram models across four canonical PoS phenomena. Findings indicate that ANN-based models can achieve above-chance generalization from limited input (10 million words), but their learning efficiency lags behind that of children as input scale increases. The study also finds that cognitively motivated inductive biases substantially improve performance, though the full results are not detailed in the abstract. The work addresses inconsistencies in previous studies, which focused on individual phenomena and used varied evaluation protocols, by providing a unified framework. The authors aim to clarify whether prior findings generalize across phenomena and learning conditions. The paper is available on arXiv under the identifier 2602.09992.

Key facts

  • The study introduces PoSHBench, a unified benchmark for evaluating the Poverty of the Stimulus Hypothesis in neural language models.
  • The benchmark covers four canonical PoS phenomena.
  • Transformer, LSTM, and n-gram models were trained and evaluated.
  • ANN-based models achieved above-chance generalization from 10 million words of input.
  • Learning efficiency of ANNs is less than that of children as input scale grows.
  • Cognitively motivated inductive biases substantially improve performance.
  • The study addresses inconsistencies in previous PoSH evaluations.
  • The paper is published on arXiv with identifier 2602.09992.

Entities

Institutions

  • arXiv

Sources