LLM Persona Panel Study Reveals Imprecise Treatment-Response Estimates in Strategic Games
There's a new preprint on arXiv, identified as 2608.00979, that looks into how large language models (LLMs) can be used as synthetic research subjects. The study focuses on a specific group of sixteen lightweight GPT-4.1 models tested in strategic games. They followed certain preregistered standards in three out of four repeated-game scenarios but missed the lower reference threshold by just 0.011. They found that variations in prompts were significant, with median shares ranging from 63% to 71% under Jeffreys alpha=0.5 and 47% to 53% under alpha=1. Additionally, estimates for finite opportunities were between 85% and 96%. The findings highlight how prompt variation and uncertainty modeling are crucial in LLM research.
Key facts
- Study uses sixteen lightweight persona-conditioned GPT-4.1 configurations
- Panel met preregistered criteria in three of four repeated-game cells
- Sole miss was 0.011 below lower reference bound
- Between-prompt shares varied from 47%-96% depending on assumptions
- Jeffreys alpha=0.5 gave 63%-71% median shares; alpha=1 gave 47%-53%
- Finite-opportunity plug-in estimates were 85%-96%
- Aggregate continuation-probability contrasts were +0.083 and +0.078
- Conservative simultaneous 95% intervals: [-0.171, +0.330] and [-0.181, +0.330]
Entities
Institutions
- arXiv