ARTFEED — Contemporary Art Intelligence

LongEarth-Bench: New Benchmark for Long-Horizon Earth Observation Reasoning

ai-technology · 2026-08-15

The introduction of LongEarth-Bench marks a significant advancement in vision-language models for long-horizon Earth observation reasoning. This benchmark features around 120,000 question-answering instances sourced from 117,000 distinct images, with an average of 15.14 frames per sequence, reaching a maximum of 30 frames. It encompasses 12 tasks, including evolution summarization, spatial reasoning, anomaly detection, and logical prediction. A subset of 30,000 samples offers structured reasoning paths that connect crucial frames and altered areas to the final answers. The LongEarth model is created through supervised fine-tuning, utilizing explicit sequence identifiers and structured chain-of-thought supervision. This research addresses the shortcomings of current remote sensing vision-language models that primarily focus on single images, pairs, or brief sequences. The paper can be found on arXiv with the identifier 2608.13344.

Key facts

  • LongEarth-Bench contains approximately 120k question-answering samples.
  • The benchmark is derived from 117k unique images.
  • Sequences average 15.14 frames and extend to 30 frames.
  • It covers 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction.
  • A 30k-sample subset provides structured reasoning traces.
  • The LongEarth model is developed through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision.
  • The paper is available on arXiv with ID 2608.13344.
  • The work targets long-horizon Earth observation reasoning, which requires organizing multi-stage geographic evolution, localizing spatial changes, detecting temporal anomalies, and inferring future from extended image sequences.

Entities

Sources