PatientAgentBench: Benchmarking Patient-Facing Health AI Agents
A new benchmark framework called PatientAgentBench evaluates foundation models acting as patient-facing health AI agents. Unlike existing benchmarks that focus on isolated medical knowledge or clinician-facing tasks, PatientAgentBench assesses agents that converse with simulated patients, reason about health records, and use healthcare tools in a sandbox environment. The evaluation uses an LLM-as-a-Jury across six dimensions with over a hundred conversation-agnostic, clinician-grounded criteria. Validation with licensed clinicians showed 79-93% adjacent agreement between the jury and expert raters, comparable to or exceeding clinician inter-rater agreement. The framework aims to address risks in primary care, particularly diagnostic errors and unsafe care.
Key facts
- PatientAgentBench is a benchmark for patient-facing health AI agents.
- It evaluates foundation models wrapped in an agent with healthcare tools.
- Agents converse with simulated patients and reason about health records.
- Scoring uses an LLM-as-a-Jury across six dimensions.
- Over a hundred clinician-grounded criteria are used.
- Licensed clinicians validated the benchmark with 79-93% adjacent agreement.
- Agreement is on par with or exceeds clinician inter-rater agreement.
- The framework targets primary care risks like diagnostic errors.
Entities
—