First Study on Annotator Performance in Drone, Aviation, and Satellite Imagery for Disaster Damage Assessment
A recent study, identified as arXiv:2608.14942, marks the inaugural empirical analysis of the performance of annotators and reviewers in assessing building damage from multi-source remotely sensed images following disasters. This research scrutinizes human labeling from drone, crewed aviation, and satellite perspectives, highlighting a significant gap in existing methodologies that typically depend on single-source imagery. The dataset encompasses 9 disasters, with 20,041 buildings labeled from drone images, 20,695 from crewed aviation, and 33,392 from satellite images, contributed by 187 annotators. Labels underwent two quality-control processes: an initial review by a single reviewer and a subsequent consensus-committee review. Findings challenge existing beliefs about the reliability of imagery sources, indicating considerable variation in performance across them. The implications of this study are crucial for optimizing human resource allocation in managing extensive multi-source aerial datasets, especially in disaster response scenarios. The paper can be accessed on arXiv using the identifier 2608.14942.
Key facts
- First empirical study of annotator and reviewer performance across drone, crewed aviation, and satellite imagery.
- Dataset includes 9 disasters with 20,041 buildings in drone, 20,695 in crewed aviation, and 33,392 in satellite imagery.
- Labels provided by 187 annotators.
- Two quality-control stages: single-reviewer pass and consensus-committee review.
- Existing aerial imagery datasets rely predominantly on single-source imagery.
- No established state of practice for allocating human labor in multi-source aerial datasets.
- Findings raise questions about performance across imagery sources.
- Paper available on arXiv:2608.14942.
Entities
Institutions
- arXiv