ARTFEED — Contemporary Art Intelligence

Clinician Preferences Fail as Proxy for LLM Clinical Safety

ai-technology · 2026-08-06

A recent investigation conducted by the Massive Open Online Validation and Evaluation (MOOVE) platform indicates that the preferences of clinicians in pairwise comparisons do not reliably reflect the clinical safety of large language models (LLMs). The study examined 26,804 pairwise evaluations from 13 LLMs, contributed by over 736 clinicians from more than 28 countries. It revealed that models that score well in pairwise preferences can still demonstrate significant clinically relevant failures (scores of -1 or lower) in areas such as Harmlessness and Accuracy. These failures vary by specialty, leading to specific 'no-go zones' for clinical application. Available on arXiv (ID 2608.02617), this clinician-led research emphasizes the necessity for more comprehensive evaluation methods beyond mere preference rankings.

Key facts

  • Study evaluates whether clinician pairwise preferences reliably signal clinical safety in LLMs.
  • Data from MOOVE platform: 26,804 pairwise judgments, 13 LLMs, 736+ clinicians, 28+ countries.
  • Clinicians used a discrete [-2, +2] scale; negative values indicate unsafe or misleading content.
  • Models ranking high under pairwise preference can still have substantial rates of clinically meaningful failures (≤ -1).
  • Failures are unevenly distributed across specialties, creating domain-specific 'no-go zones'.
  • Study published on arXiv with ID 2608.02617.
  • Research conducted by a clinician-led team.
  • Findings suggest pairwise preference is a poor proxy for safety-critical performance.

Entities

Institutions

  • MOOVE (Massive Open Online Validation and Evaluation)
  • arXiv

Sources