ARTFEED — Contemporary Art Intelligence

CosmosAlign: Adapting World Foundation Model for Traffic Video Forecasting

ai-technology · 2026-08-11

A novel framework named CosmosAlign has been unveiled for the purpose of generative traffic video forecasting. This system generates long-term, temporally consistent future traffic scene videos based on brief observation histories and textual inputs. It utilizes the pretrained Cosmos3-Nano world foundation model. The methodology highlights that effectively adapting large pretrained world models for forecasting tasks relies more on distribution alignment than on enhancing model capacity. To achieve this, a two-stage LoRA adaptation strategy is introduced: the first stage aligns the conditioning-mode distribution with the target forecasting task, while the second stage synchronizes training captions with the model's structured prompting interface via an LLM-based re-captioning process. Prediction quality during inference is enhanced through a training-free consensus-based sampling method. The research can be found on arXiv under the identifier 2608.07693.

Key facts

  • CosmosAlign is a generative traffic video forecasting framework.
  • It is built upon the pretrained Cosmos3-Nano world foundation model.
  • The approach emphasizes distribution alignment over model capacity.
  • A two-stage LoRA adaptation strategy is proposed.
  • The first stage aligns conditioning-mode distribution with the forecasting task.
  • The second stage aligns training captions with the model's native structured prompting interface.
  • An LLM-based re-captioning pipeline is used for caption alignment.
  • A training-free consensus-based procedure improves inference quality.
  • The paper is available on arXiv with identifier 2608.07693.

Entities

Institutions

  • arXiv

Sources