ARTFEED — Contemporary Art Intelligence

CARE-Bench: New Benchmark Evaluates Patient-Facing LLM Triage Accuracy

ai-technology · 2026-08-06

CARE-Bench has been launched by researchers as a benchmark aimed at assessing the safety and precision of medical large language models (LLMs) and agents that respond to symptom inquiries prior to clinician interaction. This benchmark emphasizes a crucial safety aspect: determining the next steps a user should take, structured as a four-label current-action task per turn. It includes 500 cases and 1,059 patient-disclosure prefixes derived from medical dialogues, consultations, and follow-up questions. The study assessed 11 models across 269 held-out rounds, utilizing both unprompted and minimally prompted open-ended methods, with a fixed GPT-5.5 mapper categorizing responses into the four-label action framework. Findings indicate that unprompted macro-F1 scores are low, between 31.2 and 50.4. Performance improves for 10 out of 11 models when prompted, with scores ranging from 46.9 to 63.4, yet significant threshold errors remain. Notably, prompted models frequently suggest care before acquiring necessary clarifications, revealing a tendency for hasty actions. Details of the benchmark are available in the arXiv paper 2608.03731, recently submitted.

Key facts

  • CARE-Bench is a source-grounded benchmark for patient-facing LLM triage.
  • It evaluates sequential triage as a four-label per-turn current-action task.
  • The benchmark contains 500 cases and 1,059 patient-disclosure prefixes.
  • Sources include medical dialogue, consultation, and follow-up-question data.
  • 11 models were evaluated on 269 held-out rounds.
  • Unprompted macro-F1 ranged from 31.2 to 50.4.
  • Prompting improved 10 of 11 models, with macro-F1 from 46.9 to 63.4.
  • Prompted models often recommend care before obtaining needed clarification.

Entities

Institutions

  • arXiv

Sources