SABRE: Scalable Automated Benchmarking for Vision-Language Models
A recent research article presents SABRE, an automated and scalable framework designed for creating stress tests for vision-language models (VLMs). Released on arXiv (2608.07435), this paper tackles the gap in benchmark development amid the swift advancements in VLMs. SABRE transforms a Test Primer—a Markdown Task Design with Data Schema—into organized specifications, which include generated or modified images and question-answer sets. An automated filtering process eliminates candidates solvable by a Filtering VLM, while human evaluation ensures the accuracy of candidates and aids in correcting annotations and localized image adjustments. The authors implement SABRE-Prior to assess if VLMs prioritize visual evidence over learned world priors. The benchmark comprises 600 images and 1,000 questions across Context, Texture, and Attribute categories.
Key facts
- SABRE is a scalable, automated pipeline for benchmarking VLMs under stress.
- It converts a Test Primer into structured specifications, images, and QA pairs.
- Automated filtering removes candidates solved by a Filtering VLM.
- Human review verifies candidate validity and supports annotation correction.
- SABRE-Prior tests whether VLMs follow visual evidence over world priors.
- The benchmark includes 600 images and 1,000 questions.
- It spans Context, Texture, and Attribute stress types.
- Paper available on arXiv (2608.07435).
Entities
Institutions
- arXiv