LingT2I: New Benchmark Exposes Cross-Lingual Gaps in Text-to-Image AI
A new research paper introduces LingT2I, a benchmark designed to evaluate cross-lingual consistency in text-to-image (T2I) generation models. The benchmark covers 10 widely used languages and includes 33,000 prompts, aiming to address the gap in research that has largely focused on English-only settings. The study reveals linguistic inequality and language-dependent trade-offs across evaluation dimensions, as well as systematic impacts of linguistic factors and cultural contexts on model outputs. The paper provides a foundation for studying cross-lingual behavior in T2I generation, highlighting the need for more inclusive AI systems. The research is available on arXiv under the identifier 2608.11002.
Key facts
- LingT2I is a benchmark for cross-lingual text-to-image generation.
- It covers 10 widely used languages.
- The benchmark includes 33K prompts.
- The study uncovers linguistic inequality and language-dependent trade-offs.
- Language and cultural contexts systematically impact model outputs.
- The research addresses the English-only focus of previous T2I studies.
- The paper is available on arXiv (2608.11002).
- The benchmark aims to facilitate cross-lingual T2I research.
Entities
Institutions
- arXiv