Base vs Post-Trained LLMs: Emulation vs Estimation in Opinion Simulation
A recent paper on arXiv (2608.03044) explores the reasons behind the inconsistent outcomes of large language models when simulating human opinions. The researchers distinguish between two tasks: emulation, where models create individual responses that combine into a population distribution, and estimation, where models aim to predict the population distribution directly. Their evaluation of six base and post-trained models using the Pew American Trends Panel reveals that base models perform better in emulation, aligning more closely with human responses and maintaining demographic integrity. Conversely, post-trained models excel in estimation, providing more precise predictions. The findings indicate that choosing the appropriate model for simulating human opinions should align with the specific task, as mixing these tasks may clarify previous discrepancies. This paper is accessible on arXiv and was presented as a cross-type submission.
Key facts
- Paper arXiv:2608.03044
- Announce Type: cross
- Study evaluates six matched base and post-trained models
- Uses Pew American Trends Panel data
- Base models are stronger emulators
- Post-trained models are stronger estimators
- Conflating emulation and estimation may explain conflicting results
- Model selection should depend on the task
Entities
Institutions
- arXiv
- Pew American Trends Panel