New Benchmark M²BIND Reveals Language-Dependent Fragility in Vision-Language Models
A recent study, titled 'Vision-Language Models are Fragile Multilingual Associators,' presents M²BIND, a benchmark aimed at assessing the stability of associations between visual elements and textual attributes in vision-language models (VLMs) when the input language varies. Published on arXiv, the research measures binding both extrinsically via task performance metrics and intrinsically through causal interventions. Results reveal that binding is not invariant across languages; significant binding collapse occurs in cross-family and cross-script scenarios, with internal computations shifting to later layers and diminishing in causal strength. Associations are maintained more effectively among closely related languages. The authors assert that VLMs used in multilingual contexts may not uphold the same quality of associations found in monolingual assessments. The paper falls under Computer Science > Computation and Language and can be accessed via the arXiv link.
Key facts
- The paper introduces M²BIND, a benchmark for evaluating multilingual binding in vision-language models.
- Binding refers to the association of visual entities with textual attributes.
- The benchmark varies the language of context and query across multiple languages.
- Evaluation includes both extrinsic task performance metrics and intrinsic causal interventions.
- Findings show binding is not language-invariant, with cross-family and cross-script settings causing significant collapse.
- Internal binding computation shifts to later layers and loses causal strength in such settings.
- Closely related languages preserve associations better than distant ones.
- The study implies VLMs in multilingual deployments may have inconsistent association quality.
Entities
Institutions
- arXiv