ARTFEED — Contemporary Art Intelligence

ClinLens: Benchmark for Longitudinal Clinical AI Agents

other · 2026-07-30

A team of researchers has unveiled CLINLENS, a new benchmark consisting of 200 tasks designed to assess clinical data-science agents using five interconnected MIMIC resources. These resources include various types of clinical data such as electronic health records and test results. The benchmark employs a structured matrix that evaluates four patient-time perspectives alongside five different analytical abilities. Using a novel program-first reverse synthesis technique, it links tasks to a coded evaluation system. Notably, the leading model out of 24 configurations scored 56.3% on a specific metric while achieving a flawless success rate in execution. Another agent demonstrated proficiency by completing 83 tasks.

Key facts

  • CLINLENS is a benchmark of 200 executable tasks.
  • Tasks span five linked MIMIC resources: structured EHR, notes, ECGs, chest X-rays, echocardiograms.
  • A 4x5 taxonomy crosses four patient-time scopes with five analysis capabilities.
  • Program-first reverse synthesis pairs semi-raw packages with evaluator-private reference workflows.
  • Checks required artifacts, cohort and temporal semantics, and final answer.
  • On a 126-task suite, best configuration achieved 56.3% scope-macro STRICTPASS.
  • Best configuration had 100% EXECSUCCESS.
  • A separately configured coding agent solved 83 of 126 tasks.

Entities

Sources