CosmosAlign: Adapting World Foundation Model for Traffic Video Forecasting
A novel framework named CosmosAlign has been unveiled for the purpose of generative traffic video forecasting. This system generates long-term, temporally consistent future traffic scene videos based on brief observation histories and textual inputs. It utilizes the pretrained Cosmos3-Nano world foundation model. The methodology highlights that effectively adapting large pretrained world models for forecasting tasks relies more on distribution alignment than on enhancing model capacity. To achieve this, a two-stage LoRA adaptation strategy is introduced: the first stage aligns the conditioning-mode distribution with the target forecasting task, while the second stage synchronizes training captions with the model's structured prompting interface via an LLM-based re-captioning process. Prediction quality during inference is enhanced through a training-free consensus-based sampling method. The research can be found on arXiv under the identifier 2608.07693.
Key facts
- CosmosAlign is a generative traffic video forecasting framework.
- It is built upon the pretrained Cosmos3-Nano world foundation model.
- The approach emphasizes distribution alignment over model capacity.
- A two-stage LoRA adaptation strategy is proposed.
- The first stage aligns conditioning-mode distribution with the forecasting task.
- The second stage aligns training captions with the model's native structured prompting interface.
- An LLM-based re-captioning pipeline is used for caption alignment.
- A training-free consensus-based procedure improves inference quality.
- The paper is available on arXiv with identifier 2608.07693.
Entities
Institutions
- arXiv