ARTFEED — Contemporary Art Intelligence

GeoBenchLLM: New Benchmark Evaluates LLMs on Geo-Spatial and Temporal Tasks

ai-technology · 2026-08-10

Researchers have introduced GeoBenchLLM, a comprehensive benchmark designed to evaluate large language models (LLMs) on geo-related tasks, addressing the lack of standardized evaluation in this domain. The benchmark leverages twelve publicly available datasets covering diverse geo-related tasks and domains, and assesses a set of LLMs on geo-spatial and temporal understanding. The study reveals that reasoning ability and model size significantly impact overall performance. GeoBenchLLM is publicly available, providing a resource for future research. The work is detailed in a paper submitted to arXiv under the Computer Science > Artificial Intelligence category, with the identifier 2608.07411. The benchmark aims to probe LLMs' generalization capabilities in the context of geodata, which have often been studied in a homogeneous setting, limiting insights. By offering a diverse set of tasks, GeoBenchLLM enables a more nuanced evaluation of LLMs' geo-spatial and temporal reasoning. The findings underscore the importance of reasoning and scale in improving performance on geo-related tasks. The dataset and code are accessible via the provided URL, facilitating further exploration and benchmarking by the research community.

Key facts

  • GeoBenchLLM is a benchmark for evaluating LLMs on geo-related tasks.
  • It uses twelve publicly available datasets from diverse geo-related tasks and domains.
  • The benchmark evaluates LLMs on geo-spatial and temporal understanding.
  • Results show that reasoning and size have a strong impact on overall performance.
  • GeoBenchLLM is publicly available.
  • The paper is submitted to arXiv under Computer Science > Artificial Intelligence.
  • The arXiv identifier is 2608.07411.
  • The study addresses the limitation of homogeneous settings in prior LLM studies on geodata.

Entities

Institutions

  • arXiv

Sources