DreamFly: New Diffusion-Based Framework for Aerial Vision-Language Navigation
Researchers have introduced DreamFly, a novel diffusion-based framework for aerial vision-language navigation (VLN), detailed in a paper published on arXiv (2608.12308). The framework addresses key challenges in aerial navigation, including limited historical context, short planning horizons, and unreliable termination detection. DreamFly builds on the Dream-VLA model and incorporates a causally aligned historical memory that uses only past observations to augment current visual representations, preventing future information leakage. Navigation is formulated as receding-horizon diffusion planning, where the policy predicts a K-step action chunk but executes only the first action before replanning. This approach enhances temporal reasoning and decision-making under partial observability. The paper is categorized as a cross-type announcement and is available at the provided arXiv URL.
Key facts
- DreamFly is a diffusion-based aerial VLN framework built on Dream-VLA.
- It introduces causally aligned historical memory to prevent future information leakage.
- Navigation is formulated as receding-horizon diffusion planning.
- The policy predicts a K-step action chunk but executes only the first action before replanning.
- The paper is published on arXiv with ID 2608.12308.
- The announcement type is 'cross'.
- The framework addresses challenges in aerial VLN: limited historical context, short planning horizons, and unreliable implicit termination.
- The paper is available at https://arxiv.org/abs/2608.12308.
Entities
Institutions
- arXiv