ORCA: Training-Free Token Compression for 3D CT Scans
A novel technique named ORCA (ORgan-Centroid Aggregation) has been developed to compress visual tokens in 3D CT scans specifically for vision-language models. This method tackles the issue of extensive token sequences produced from 3D CT volumes, which can vary from thousands to tens of thousands per scan. Conventional grid average pooling often merges different anatomical features, lesions, and air into a single token, resulting in the loss of essential details. ORCA utilizes organ guidance to combine neighboring tokens and applies a sinusoidal encoding of each region's centroid to maintain spatial organization. This plug-and-play, training-free approach yields a customizable token set without necessitating changes to the model or text queries. The evaluation of this method was conducted on two datasets, CT-RATE and Merlin, across five encoders and two task types. The research paper can be found on arXiv with the identifier 2608.00345.
Key facts
- ORCA is a token compressor for 3D CT scans.
- It merges adjacent tokens with organ guidance.
- Adds sinusoidal encoding of region centroids to preserve spatial layout.
- Training-free and plug-and-play.
- Produces adjustable token set without model change or text query.
- Evaluated on CT-RATE and Merlin datasets.
- Tested across five encoders.
- Evaluation spans two task types.
- Paper available on arXiv:2608.00345.
Entities
Institutions
- arXiv