TangPoetryBench: New Benchmark for Evaluating Poetry-to-Image Generation
Researchers have introduced TangPoetryBench, a multi-dimensional benchmark designed to evaluate text-to-image (T2I) models on their ability to illustrate classical Chinese Tang poetry. The benchmark comprises 1,280 images generated from 320 poems using four state-of-the-art T2I models, accompanied by quality-controlled human annotations across ten dimensions. The study addresses the limitations of existing metrics such as CLIPScore, BLIPScore, and VQAScore, which only measure literal text-image correspondence and fail to capture the nuanced requirements of poetry illustration, including visual soundness, fidelity to imagery and scene, cultural and stylistic appropriateness, absence of spurious text, and emotional truth. The analysis reveals shared and model-specific strengths and weaknesses, providing insights into the current capabilities of T2I models in rendering literary and cultural content. The paper is available on arXiv under the identifier 2608.11452.
Key facts
- TangPoetryBench includes 1,280 images generated from 320 classical Chinese Tang poems.
- Four state-of-the-art T2I models were used to generate the images.
- Human annotations cover ten dimensions of evaluation.
- Existing metrics like CLIPScore, BLIPScore, and VQAScore are inadequate for poetry illustration.
- The benchmark assesses visual soundness, fidelity, cultural aptness, and emotional truth.
- The study identifies shared and model-specific strengths and weaknesses.
- The paper is available on arXiv (2608.11452).
- The task of poetry illustration is many-sided and includes implicit emotion.
Entities
Institutions
- arXiv