MedPIC-Bench: Benchmarking Patient-Specific Medication-Safety Reasoning in LLMs
A new benchmark, MedPIC-Bench, has been introduced to evaluate the ability of large language models (LLMs) to apply medication-safety rules based on patient-specific information. The benchmark, detailed in a paper on arXiv (2608.03028), addresses a critical gap in medical AI evaluation: existing tests often use isolated, fixed scenarios, allowing models to answer correctly by recalling drug-risk associations without demonstrating that they considered patient information to determine whether a rule applies. MedPIC-Bench consists of 467 questions, each annotated along six clinical and reasoning dimensions, and combines guideline-following questions with paired counterfactual questions. In counterfactual scenarios, a controlled change in patient information alters whether a rule applies, testing the model's sensitivity to patient context. The study evaluated 28 medical-specific, general, and proprietary LLMs, finding that every model performed worse on counterfactual questions, with mean accuracy dropping from 63.6% to 45.1%. This significant performance gap highlights a common weakness in LLMs' reasoning about patient-specific conditions. The benchmark aims to promote more robust and safe AI applications in healthcare by encouraging models to genuinely reason about patient information rather than relying on superficial associations.
Key facts
- MedPIC-Bench is a new benchmark for patient-specific medication-safety reasoning.
- It includes 467 questions annotated along six clinical and reasoning dimensions.
- The benchmark combines guideline-following questions with paired counterfactual questions.
- 28 LLMs (medical-specific, general, and proprietary) were evaluated.
- All models performed worse on counterfactual questions.
- Mean accuracy fell from 63.6% to 45.1% on counterfactual questions.
- The benchmark addresses a gap in existing medical evaluations.
- The paper is available on arXiv with ID 2608.03028.
Entities
Institutions
- arXiv