ARTFEED — Contemporary Art Intelligence

AI Models Confabulate Medical Diagnoses Based on Patient Demographics

ai-technology · 2026-07-30

A recent study published on arXiv (2607.26886) indicates that advanced vision-language models, such as Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro, do not refrain from providing diagnoses when prompted with a medical image that isn’t actually provided. Instead, they generate fabricated diagnoses influenced by the demographic details of the patient. This effect is evident across various imaging types, including chest X-rays and brain MRIs. For instance, a 65-year-old white male inquiring about a skin mole is frequently diagnosed with Melanoma, whereas a 32-year-old Black female asking about her chest X-ray is often diagnosed with Sarcoidosis, justified by "suspected, based on demographics and classic pattern." GPT-5.4 demonstrates a wider range, particularly fabricating Sarcoidosis for young Black patients in chest X-rays. The findings underscore how these models reinforce demographic biases in medical AI.

Key facts

  • Frontier vision-language models confabulate diagnoses when no image is provided.
  • Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro were tested.
  • Demographic descriptors systematically shift the confabulated diagnosis.
  • Claude diagnoses Melanoma for a 65-year-old white man with a skin mole.
  • Claude diagnoses Sarcoidosis for a 32-year-old Black woman with a chest X-ray.
  • GPT-5.4 fabricates diagnoses across all demographic cells tested.
  • The study covers chest X-ray, brain MRI, and dermatology.
  • The paper is available on arXiv (2607.26886).

Entities

Institutions

  • arXiv

Sources