Anatomy Contextualized Adaptation Enhances CT Foundation Models
A new lightweight framework called Anatomy Contextualized Adaptation (ACA) improves CT vision-language foundation models by aligning anatomy-level visual features with text while preserving global context. ACA adapts frozen whole-volume model representations using TotalSegmentator to decompose CT scans into anatomy-level embeddings, refined via a transformer that captures cross-anatomy relationships. This approach avoids training from scratch, reducing computational costs. The method addresses limitations of existing fine-grained pre-training that discards global context and requires full retraining. The paper is published on arXiv under ID 2607.27154.
Key facts
- ACA is a lightweight framework for adapting frozen CT foundation models.
- It uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings.
- A transformer refines embeddings to capture cross-anatomy relationships.
- Aligns per-anatomy and scan-level text for vision-language alignment.
- Avoids training from scratch, reducing computational expense.
- Addresses dilution of fine-grained anatomical signals in whole-volume models.
- Preserves global context that fine-grained approaches discard.
- Published on arXiv with ID 2607.27154.
Entities
Institutions
- arXiv