MissionBench: Benchmarking Aerial MLLM Agents on Long-Horizon Tasks
MissionBench has been unveiled by researchers as a benchmark designed to assess multimodal large language models (MLLMs) functioning as embodied agents within aerial 3D settings. This benchmark features 120 missions spread across five simulated environments and four categories of tasks, challenging agents to independently plan, navigate, and report results solely based on egocentric observations and their action history, without any fine-tuning specific to aerial tasks. An evaluation of 22 MLLMs, both open and closed-source, showed that even the top-performing model managed to complete less than 35% of the missions, while human performance reached 84.4%, underscoring the complexity of multi-step embodied tasks. Notably, scaling up models led to improved performance, suggesting that larger general-purpose models exhibit enhanced zero-shot embodied abilities.
Key facts
- MissionBench is a benchmark for mission-level evaluation of MLLMs in aerial 3D environments.
- It includes 120 missions across five simulated 3D environments and four task families.
- Agents must plan, navigate, and report outcomes using egocentric observations and action history.
- No aerial-specific fine-tuning is allowed.
- 22 open- and closed-source MLLMs were tested.
- The strongest model succeeded on fewer than 35% of missions.
- Human performance was 84.4%.
- Scaling improved performance across model families.
Entities
—