Multimodal Models Show Reduced Semantic Diversity in Image Generation
A new study on arXiv (2608.00410) investigates how multimodal models handle polysemous words—terms with multiple meanings like 'bank' or 'palm'. The research, which analyzed 17 text-to-image and 15 text-generation models, reveals a significant gap: when given a polysemous word without context, image generation models produce far fewer distinct senses than text models. Specifically, the normalized entropy for images was 0.10, compared to 0.25 for text, and both were much lower than human imagination at 0.47. The study also found that when models are asked to predict the distribution of meanings they would generate, they can anticipate their own output distribution, suggesting a potential method for calibration. This research highlights a limitation in current AI systems' ability to capture the full semantic richness of language, which has implications for creative applications and human-AI interaction.
Key facts
- The study is published on arXiv with ID 2608.00410.
- It examines 17 text-to-image and 15 text-generation models.
- Polysemous words like 'bank' and 'palm' were used as stimuli.
- Image generation models showed normalized entropy of 0.10 for senses produced.
- Text generation models had normalized entropy of 0.25.
- Human imagination had normalized entropy of 0.47.
- Models can predict their own output distributions for polysemous words.
- The study identifies a 'multimodal gap' in semantic diversity.
Entities
Institutions
- arXiv