ARTFEED — Contemporary Art Intelligence

Cheap Probes Could Replace Expensive Training for 3D-CT Vision-Language Models

other · 2026-07-29

A new arXiv preprint (2607.22771) proposes using cheap probes on cached embeddings to evaluate encoder and token-compression combinations for 3D CT vision-language models, avoiding costly fine-tuning of large language models. The method introduces an image-grounded probing benchmark with clinical attribute families and two validation gates: scale-sanity and probe-separability. These gates ensure attributes are well-scaled and decodable. The study compares various read-out heads in a preliminary investigation, aiming to reduce the computational burden of searching over many encoder-compression-token budget combinations.

Key facts

  • arXiv preprint 2607.22771 proposes cheap probes for 3D CT VLM evaluation
  • Probes use cached embeddings from frozen image encoders
  • Avoids fine-tuning large language models on each combination
  • Introduces image-grounded probing benchmark with clinical attribute families
  • Two validation gates: scale-sanity and probe-separability
  • Compares range of read-out heads in preliminary study
  • Aims to reduce computational cost of encoder-compression search
  • Focuses on 3D computed tomography vision-language models

Entities

Institutions

  • arXiv

Sources