Frozen Brain-MRI Foundation Models Are Site Fingerprints
A recent preprint on arXiv (2608.10295) indicates that frozen foundation-model embeddings for brain MRI predominantly represent the acquisition site rather than anatomical details. The research evaluated the actual content captured by these standard representations, revealing that across two separate cohorts (ABIDE-I and ABIDE-II), three types of frozen 3D encoders (brain-pretrained, CT-pretrained, and randomly initialized), and all network depths, the acquisition site can be linearly decoded with approximately 0.9 balanced accuracy in deeper layers. This level of decodability surpasses that of any clinical or demographic factors (sex, age, autism diagnosis) at every layer. The intrinsic nature of this effect is highlighted by a randomly initialized encoder achieving ~0.9 site classification across all three architecture families (Swin, ViT, ResNet) and both cohorts. Additionally, raw downsampled images can yield ~0.95 site decodability without an encoder, suggesting that the fingerprint reflects low-level image statistics. These results challenge the notion that frozen embeddings convey significant anatomical information, implying they may unintentionally act as site identifiers, potentially complicating multi-site studies and transfer learning in medical imaging.
Key facts
- Preprint arXiv:2608.10295
- Frozen foundation-model embeddings for brain MRI encode acquisition site
- Two cohorts: ABIDE-I and ABIDE-II
- Three encoders: brain-pretrained, CT-pretrained, randomly initialized
- Site decodable at ~0.9 balanced accuracy at deep layers
- Randomly initialized encoder achieves ~0.9 site classification
- Site decodable at ~0.95 from raw downsampled images
- Three architecture families: Swin, ViT, ResNet
Entities
Institutions
- arXiv