UAVQA-Bench: A New Benchmark for MLLM-Based UAV Image Understanding
Researchers have created a new evaluation tool named UAVQA-Bench to measure how well Multimodal Large Language Models (MLLMs) can understand and reason with UAV aerial images. This tool features 1,500 human-annotated question-answer pairs sourced from 13 publicly accessible UAV datasets. It covers six capability dimensions and includes 16 different tasks in both multiple-choice and visual grounding formats. The study tested various MLLMs, both open-source and closed-source, as well as agent-based systems, identifying three main failure points: mismatches in domain-toolsets, unchecked error propagation, and static reasoning issues. The findings aim to standardize evaluations and improve assessments related to UAV understanding. You can find the paper on arXiv using the identifier 2608.11738.
Key facts
- UAVQA-Bench is a new benchmark for MLLM-based UAV image understanding and reasoning.
- It includes 1,500 human-annotated QA pairs from 13 public UAV datasets.
- The benchmark covers 6 capability dimensions and 16 tasks.
- Tasks are in multiple-choice and visual grounding formats.
- Systematic evaluation of open-source and closed-source MLLMs and agent-based systems was conducted.
- Three key failure modes were identified: domain-toolset mismatch, unchecked error propagation, and static reasoning.
- The paper is available on arXiv with the identifier 2608.11738.
Entities
—