EyExIn: Deep Expert Injection Anchors Retinal VLMs with Domain Knowledge
So, a group of researchers has come up with this new framework called EyExIn to address some major issues with Large Vision Language Models (LVLMs) that are used for diagnosing eye diseases. They identified two main problems: the Perception Gap, where the models miss tiny signs of disease like microaneurysms, and the Reasoning Gap, where biases in language can overshadow visual cues. EyExIn uses a method called Deep Expert Injection and an Expert-Aware Dual-Stream encoding system. This separates the visual data into a general stream for anatomical details and a specialized one for expert knowledge. You can check out this research on arXiv, under the reference 2603.07131v4.
Key facts
- EyExIn addresses the Perception Gap and Reasoning Gap in retinal LVLMs.
- Perception Gap: general visual encoders fail on fine pathological cues like microaneurysms.
- Reasoning Gap: language priors override visual evidence in deeper transformer layers.
- EyExIn uses a Deep Expert Injection mechanism.
- Architecture: Expert-Aware Dual-Stream encoding.
- General stream handles anatomical context.
- Specialized expert stream handles domain-specific knowledge.
- Published on arXiv with ID 2603.07131v4.
Entities
Institutions
- arXiv