MedUP: A New Medical Vision-Language Model Unifying Perception and Understanding
A team of researchers has unveiled MedUP, an innovative Medical Vision-Language Model (Med-VLM) that integrates perception and comprehension within a common token framework. This model tackles issues related to accurate visual perception, segmentation, and grounding that current Med-VLMs encounter. Conventional methods typically express regions as coordinate strings or depend on external modules, leading to gaps in region-language alignment. The standout feature of MedUP is UniMedTok, a region tokenizer that transforms masks into discrete tokens within the LLM vocabulary, enabling seamless integration of mask tokens with text. The researchers compiled UniMed-Train, a substantial dataset of 1.84 million instances encompassing text-guided segmentation, region-grounded understanding, medical VQA, and chain-of-thought (CoT) segmentation. They also developed UniMed-Bench for comprehensive evaluation. Rigorous testing shows that MedUP surpasses native, agentic, and dual-decoder Med-VLMs across all metrics. The paper can be found on arXiv with the identifier 2608.10635.
Key facts
- MedUP is a new Medical Vision-Language Model that unifies perception and understanding in a shared token space.
- The model uses UniMedTok, a region tokenizer that encodes masks as discrete tokens in the LLM vocabulary.
- UniMed-Train is a 1.84M-instance corpus for training, covering text-guided segmentation, region-grounded understanding, medical VQA, and CoT-based segmentation.
- UniMed-Bench is introduced for unified evaluation of Med-VLMs.
- MedUP outperforms native, agentic, and dual-decoder Med-VLMs across all benchmarks.
- The paper is available on arXiv with identifier 2608.10635.
- The research addresses representation gaps in region-language alignment in existing Med-VLMs.
- The model natively interleaves mask tokens with text, avoiding external modules.
Entities
Institutions
- arXiv