Benchmarking 3D CT Foundation Models Reveals No Universal State-of-the-Art
A recent preprint on arXiv (2608.05960) assesses ten frozen 3D CT foundation models across three cohorts of thoracic CT scans, including a previously unseen internal clinical dataset. The research employs k-nearest neighbors, zero-shot prompting, and linear probing for evaluation. The findings indicate that there is no single state-of-the-art model, as rankings differ based on the evaluation context. Generally, models that integrate fine-grained image tokenization with vision-language alignment yield the best results. However, a lightweight supervised encoder also shows strong competitiveness, suggesting that explicit labels can effectively replace scale. Performance is primarily limited by a physical bottleneck rather than the model's architecture, impacting routine CT interpretation, particularly regarding incidental findings.
Key facts
- Ten frozen 3D CT encoders were benchmarked.
- Three cohorts of thoracic CT scans were used, including an unseen internal clinical dataset.
- Evaluation methods included k-nearest neighbors, zero-shot prompting, and linear probing.
- No universal state-of-the-art model was found.
- Rankings fluctuated significantly depending on the evaluation context.
- Models combining fine-grained image tokenization with vision-language alignment performed best overall.
- A lightweight supervised encoder remained highly competitive.
- The primary determinant of performance was a physical bottleneck, not model architecture.
Entities
Institutions
- arXiv