LENS: Decomposing Vision-Language Model Representations
A recent study presents LENS (Local Explanation of Neighborhood Subspaces), a technique that utilizes a Mixture of Factor Analyzers to break down activations from vision-language models (VLM) into localized low-rank Gaussian neighborhoods. This approach overcomes the drawbacks of global linear interpretability methods, which often overlook locally low-dimensional representations. When applied to LLaVA-1.5-7B and Qwen3-VL-8B, LENS uncovers unique depth-dependent fusion paths: LLaVA integrates modalities in later layers, whereas Qwen3-VL initiates mixing early, partially re-segregates, and then recombines close to the output. An automated pipeline for multimodal labeling provides semantic descriptions for these neighborhoods. The research can be found on arXiv under ID 2608.00561.
Key facts
- LENS decomposes VLM activations into local low-rank Gaussian neighborhoods
- Uses Mixture of Factor Analyzers
- Applied to LLaVA-1.5-7B and Qwen3-VL-8B
- Reveals depth-dependent fusion trajectories
- LLaVA mixes modalities at later layers
- Qwen3-VL mixes early, re-segregates, and recombines
- Automated multimodal labeling pipeline assigns semantic descriptions
- Paper available on arXiv with ID 2608.00561
Entities
Institutions
- arXiv