LLMs Fail as Synthetic Users in Survey Simulation Benchmark
A new study from arXiv (2607.26348) systematically evaluates whether large language models (LLMs) can replace human respondents in survey-based research. The researchers tested four models from two families (ranging from 8B to frontier capability) on two independent datasets: the U.S. General Social Survey and the World Values Survey. Under demographic prompting and standardized simulation protocols, no LLM outperformed the strongest non-LLM baseline fitted on held-out human data. At the individual level, all models failed to replicate human response patterns. The failures were consistent across both domains, all four models, and both model families. The study provides an evaluation framework for synthetic-user systems and highlights the limitations of current LLMs in simulating human survey responses.
Key facts
- Four LLMs from two families tested (8B to frontier capability)
- Two datasets: General Social Survey (U.S.) and World Values Survey (cross-cultural)
- No LLM beat the strongest non-LLM baseline at individual level
- Failures replicated across both domains, all models, and both families
- Study proposes evaluation framework for synthetic-user systems
- arXiv paper ID: 2607.26348
- Published as arXiv preprint
- Focus on demographic prompting and survey-simulation protocols
Entities
Institutions
- arXiv
- General Social Survey
- World Values Survey
Locations
- United States