ARTFEED — Contemporary Art Intelligence

WHBench: New Benchmark Exposes LLM Failures in Women's Health

ai-technology · 2026-07-27

WHBench, a novel benchmark, assesses large language models in the realm of women's health, indicating that none achieve over 75% accuracy. Comprising 47 scenarios developed by experts across 10 topics, it utilizes a rubric with 23 criteria for evaluation. The highest-performing model attained a score of 72.1%, while leading models displayed low rates of complete correctness and considerable discrepancies in harm rates. The research underscores various failure modes, including reliance on outdated guidelines, unsafe omissions, dosing mistakes, and neglect of equity issues. Additionally, inter-rater reliability is found to be moderate, and safety-weighted penalties are implemented.

Key facts

  • WHBench is a targeted evaluation suite for women's health.
  • It includes 47 expert-crafted scenarios across 10 topics.
  • 22 models were evaluated using a 23-criterion rubric.
  • No model mean performance exceeds 75%; best model scores 72.1%.
  • Top models show low fully correct rates and substantial variation in harm rates.
  • Failure modes include outdated guidelines, unsafe omissions, dosing errors, and equity blind spots.
  • Safety-weighted penalties and server-side score recalculation are used.
  • 3,102 attempted responses, 3,100 scored.

Entities

Institutions

  • arXiv

Sources