COMEX: New Benchmark and Framework for Explainable Aesthetic Image Cropping
A new study has introduced COMEX, a framework and benchmark aimed at making aesthetic image cropping more understandable. Available on arXiv (2608.07570), it reimagines the cropping task as a structured challenge that considers both composition and explanations, addressing flaws in existing methods that mainly focus on generating text after the fact. COMEX includes 33,161 quadruples, each containing an expanded image, a crop box, a category for composition, and a grounded explanation. The framework allows for concurrent learning in crop localization, understanding composition, and generating explanations. The authors suggest a two-step method, SFT+GRPO, which starts with supervised fine-tuning and follows up with group relative policy optimization to enhance the quality of explanations, significantly advancing the fields of computational aesthetics and computer vision.
Key facts
- COMEX is a new benchmark for explainable aesthetic image cropping.
- It contains 33,161 quadruples, each with an expanded image, crop box, composition category, and explanation.
- The benchmark is built through image expansion and an IO-reversal pipeline.
- The proposed framework uses a two-stage SFT+GRPO approach.
- The paper is available on arXiv with identifier 2608.07570.
- The work reformulates explainable cropping as a structured crop-composition-explanation problem.
- It enables joint learning of crop localization, composition understanding, and explanation generation.
- The approach addresses limitations of existing post-hoc explanation methods.
Entities
Institutions
- arXiv