DisasterTD: Multimodal LLM Framework for Disaster Toponym Disambiguation
Researchers propose DisasterTD, a framework that uses multimodal large language models (MLLMs) and cross-view geolocalization to disambiguate vague geographic references in social media imagery (SMI) during disasters. The system first extracts toponyms and generates candidate locations from noisy text using MLLMs, then cross-matches SMI with remote sensing imagery (RSI) and street-view imagery (SVI) to verify results. Evaluation on the Hurricane Harvey dataset shows improved geolocalization accuracy. The approach addresses a key challenge in using SMI for situational awareness and emergency response.
Key facts
- DisasterTD integrates MLLM-based semantic reasoning with cross-view geolocalization.
- MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs.
- Cross-view matching between SMI, RSI, and optionally SVI verifies and refines candidate results.
- Evaluated on the Hurricane Harvey dataset.
- SMI provides timely ground perspectives for situational awareness and emergency response.
- Geographic references in SMI are often vague or ambiguous.
- Framework aims to improve accurate geolocalization of disaster-related SMI.
- Published on arXiv with ID 2607.24856.
Entities
Institutions
- arXiv