Cheap Probes Could Replace Expensive Training for 3D-CT Vision-Language Models
A new arXiv preprint (2607.22771) proposes using cheap probes on cached embeddings to evaluate encoder and token-compression combinations for 3D CT vision-language models, avoiding costly fine-tuning of large language models. The method introduces an image-grounded probing benchmark with clinical attribute families and two validation gates: scale-sanity and probe-separability. These gates ensure attributes are well-scaled and decodable. The study compares various read-out heads in a preliminary investigation, aiming to reduce the computational burden of searching over many encoder-compression-token budget combinations.
Key facts
- arXiv preprint 2607.22771 proposes cheap probes for 3D CT VLM evaluation
- Probes use cached embeddings from frozen image encoders
- Avoids fine-tuning large language models on each combination
- Introduces image-grounded probing benchmark with clinical attribute families
- Two validation gates: scale-sanity and probe-separability
- Compares range of read-out heads in preliminary study
- Aims to reduce computational cost of encoder-compression search
- Focuses on 3D computed tomography vision-language models
Entities
Institutions
- arXiv