ARTFEED — Contemporary Art Intelligence

Benchmarking 3D CT Foundation Models Reveals No Universal State-of-the-Art

ai-technology · 2026-08-07

A recent preprint on arXiv (2608.05960) assesses ten frozen 3D CT foundation models across three cohorts of thoracic CT scans, including a previously unseen internal clinical dataset. The research employs k-nearest neighbors, zero-shot prompting, and linear probing for evaluation. The findings indicate that there is no single state-of-the-art model, as rankings differ based on the evaluation context. Generally, models that integrate fine-grained image tokenization with vision-language alignment yield the best results. However, a lightweight supervised encoder also shows strong competitiveness, suggesting that explicit labels can effectively replace scale. Performance is primarily limited by a physical bottleneck rather than the model's architecture, impacting routine CT interpretation, particularly regarding incidental findings.

Key facts

  • Ten frozen 3D CT encoders were benchmarked.
  • Three cohorts of thoracic CT scans were used, including an unseen internal clinical dataset.
  • Evaluation methods included k-nearest neighbors, zero-shot prompting, and linear probing.
  • No universal state-of-the-art model was found.
  • Rankings fluctuated significantly depending on the evaluation context.
  • Models combining fine-grained image tokenization with vision-language alignment performed best overall.
  • A lightweight supervised encoder remained highly competitive.
  • The primary determinant of performance was a physical bottleneck, not model architecture.

Entities

Institutions

  • arXiv

Sources