ClinLens: Benchmark for Longitudinal Clinical AI Agents
A team of researchers has unveiled CLINLENS, a new benchmark consisting of 200 tasks designed to assess clinical data-science agents using five interconnected MIMIC resources. These resources include various types of clinical data such as electronic health records and test results. The benchmark employs a structured matrix that evaluates four patient-time perspectives alongside five different analytical abilities. Using a novel program-first reverse synthesis technique, it links tasks to a coded evaluation system. Notably, the leading model out of 24 configurations scored 56.3% on a specific metric while achieving a flawless success rate in execution. Another agent demonstrated proficiency by completing 83 tasks.
Key facts
- CLINLENS is a benchmark of 200 executable tasks.
- Tasks span five linked MIMIC resources: structured EHR, notes, ECGs, chest X-rays, echocardiograms.
- A 4x5 taxonomy crosses four patient-time scopes with five analysis capabilities.
- Program-first reverse synthesis pairs semi-raw packages with evaluator-private reference workflows.
- Checks required artifacts, cohort and temporal semantics, and final answer.
- On a 126-task suite, best configuration achieved 56.3% scope-macro STRICTPASS.
- Best configuration had 100% EXECSUCCESS.
- A separately configured coding agent solved 83 of 126 tasks.
Entities
—