Polish Medical VQA Benchmark Reveals Vision-Language Models Underuse Visual Evidence
A new standard for Polish-language medical visual question answering (VQA) has been established, utilizing questions from the Polish Board Certification Examination aimed at licensed physicians and dentists seeking specialist credentials. This benchmark features questions that include images across various medical fields and visual categories, alongside a control set focused solely on text. An evaluation of Polish-centric, general-purpose open-weight, and commercial vision-language models reveals that the task is still quite difficult: the top-performing model reaches an accuracy of 79.0% on the complete VQA set, while only GPT-5.6 exceeds the human reference on the subset with available candidate responses; other models fall short compared to human performance. To evaluate visual grounding, researchers analyzed complete inputs against versions lacking the image, the question, or both, categorizing questions based on the significance of the image. Results show that models tend to extract more valuable information from the text than the images, indicating a tendency to underutilize visual data. The benchmark and related code can be found on arXiv under the identifier 2608.12928.
Key facts
- The benchmark is built from Polish Board Certification Examination questions for physicians and dentists.
- The benchmark includes image-containing questions and a text-only QA control set.
- The best model achieves 79.0% accuracy on the full VQA set.
- Only GPT-5.6 surpasses the approximate human reference on the subset with available candidate responses.
- All other evaluated models perform worse than humans.
- Models derive more useful information from the question text than from the image.
- The study compares complete inputs with configurations omitting the image, the question, or both.
- The benchmark is available on arXiv under identifier 2608.12928.
Entities
Institutions
- arXiv