ARTFEED — Contemporary Art Intelligence

AI Mental Health Simulations Fail Population-Level Accuracy

ai-technology · 2026-08-07

A new study from arXiv (2604.17359) reveals that large language models (LLMs) simulating psychiatric patients produce convincing individual cases but fail to replicate real-world population statistics. Researchers tested GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3, and GLM-4.7 across 120 demographic cohorts, using two framings: one as a clinician's entry and one as a self-description. They generated 28,800 responses and scored them against survey-weighted PHQ-8 anchors from NHANES microdata. Individually, 97.3% of elevated presentations met the DSM-5 gateway rule, with a violation rate of 2.68% versus a chance null of 10.4%. However, as populations, the simulations consistently inflated PHQ-8 scores by 2.8 to 5.5 points across all benchmarkable groups. Additionally, 18.2% of simulated patients screened at the treatment threshold, compared to 7.5% of real adults. The simulations also failed to preserve racial disparities: Black-White and Hispanic-White gaps were either attenuated or flattened by the models. The study highlights a critical limitation in using LLMs for mental health research and simulation, emphasizing that while they may pass individual scrutiny, they cannot accurately represent population-level epidemiological patterns.

Key facts

  • Study from arXiv:2604.17359
  • Tested GPT-4o-mini, Gemini-3-Flash, DeepSeek-V3, GLM-4.7
  • 120 demographic cohorts, two framings
  • 28,800 responses scored against PHQ-8 anchors
  • 97.3% of elevated presentations met DSM-5 gateway rule
  • Violation rate 2.68% vs chance null 10.4%
  • PHQ-8 scores inflated by 2.8 to 5.5 points
  • 18.2% simulated patients at treatment threshold vs 7.5% real adults
  • Black-White and Hispanic-White disparities not preserved

Entities

Institutions

  • arXiv

Sources