AI-Generated Synthetic Survey Data Fails to Replicate Human Insights in Silicon Valley Study
A recent investigation shared on arXiv examines the performance of synthetic survey data generated by five prominent large language models (LLMs) compared to responses from 420 human developers and coders situated in Silicon Valley. The models assessed were ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro, Claude CoWork 1.123, Gemini Advanced 2.5 Pro, and DeepSeek 3.2. The study revealed that while AI outputs appeared coherent and trustworthy, they lacked the originality and depth found in human insights. This disparity emphasizes the challenges in depending on synthetic data for organizational research aimed at uncovering innovative viewpoints.
Key facts
- Study compares synthetic data from five LLMs with human survey responses.
- Human survey involved 420 Silicon Valley coders and developers.
- LLMs tested: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, DeepSeek 3.2.
- AI-generated data was technically plausible but missed counterintuitive insights.
- Deviations from human data grouped together across all models.
- Real human data was the outlier in the analysis.
- Study published on arXiv with ID 2603.00059.
- Implications for organizational research practice are discussed.
Entities
Institutions
- arXiv
Locations
- Silicon Valley