New Scaling Law Predicts VLM Performance from LLM Textual Capabilities
A recent paper published on arXiv introduces the Capability-Driven Multimodal Scaling Law, a new framework aimed at forecasting the performance of vision-language models (VLMs) based on the directly measurable textual capabilities of their large language model (LLM) foundations. This research tackles the issue of how to appropriately choose LLM backbones for VLM development, a currently arbitrary choice. The framework derives a low-dimensional capability score from LLM textual benchmarks using PCA, modeling VLM performance in relation to this score while factoring in transfer and absorption rates for data efficiency. To test this framework, over 150 VLMs were trained on 34 LLMs across 7 model families, evaluated against more than 200 textual and 50 multimodal benchmarks. The paper can be found on arXiv with ID 2608.00013, offering a valuable approach for the AI community to predict VLM performance prior to training, thereby conserving resources and aiding in model selection.
Key facts
- Proposes Capability-Driven Multimodal Scaling Law
- Predicts VLM performance from LLM textual capabilities
- Uses PCA to extract capability score from textual benchmarks
- Trained over 150 VLMs on 34 LLMs from 7 model families
- Evaluated on more than 200 textual and 50 multimodal benchmarks
- Paper available on arXiv with ID 2608.00013
- Addresses lack of principled framework for LLM backbone selection in VLMs
- Framework includes transfer rates and absorption rates
Entities
Institutions
- arXiv