CliniCARE-Bench: New Benchmark for Clinical Audit in EHR
A new benchmark called CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR) has been developed by researchers to assess large language models (LLMs) in performing retrospective audits of clinical data using electronic health records (EHR). This benchmark includes 25 scenarios validated by clinicians, represented as 750 patient-specific cases sourced from the MIMIC-IV dataset, which features actual patient information. The systems must analyze each case within a controlled, logged environment for data retrieval, computation, and policy access, ultimately providing one of four outcomes: Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous. The latter two outcomes help differentiate between absent evidence and medical uncertainty. This benchmark is crucial for ensuring LLMs can conduct reliable investigations across diverse, longitudinal records, determining necessary evidence, retrieving structured and unstructured data, basing conclusions on verifiable evidence, and managing unresolved cases. This research is pertinent to the integration of AI in clinical environments, emphasizing the importance of reliability and accountability. The study can be found on arXiv with the identifier 2608.07796.
Key facts
- CliniCARE-Bench is a benchmark for retrospective clinical audit.
- It includes 25 clinician-validated scenarios.
- The benchmark comprises 750 patient-specific cases.
- Cases are derived from the MIMIC-IV dataset.
- Systems must return one of four verdicts: Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous.
- The benchmark uses a governed, logged tool environment for record retrieval, computation, and policy access.
- The paper is available on arXiv with identifier 2608.07796.
- The benchmark aims to evaluate LLMs' ability to conduct defensible investigations over EHR data.
Entities
Institutions
- arXiv