ARTFEED — Contemporary Art Intelligence

VTM-Nav: Cross-Episode Object-Goal Navigation with Visual-Topological Memory

other · 2026-07-27

A new framework called VTM-Nav has been introduced by researchers for object-goal navigation across episodes without requiring training. This system employs a Hierarchical Visual-Topological Memory (VTM) that retains scene information throughout different episodes without the need for retraining. VTM organizes visual memories based on basic room layouts, differentiates between in-room and remote-visible data, and keeps cues for successful approaches. When a new task arises, VTM-Nav re-localizes the agent within the established scene structure, retrieves relevant records from likely rooms, and uses memory guidance to inform current candidates. This study tackles the issue faced by training-free ObjectNav agents, which often lose previously acquired scene knowledge after each task. VTM-Nav functions with fixed navigation components and model parameters, leveraging cross-episode experiences without relying on oracle data.

Key facts

  • VTM-Nav is a training-free framework for cross-episode ObjectNav
  • Uses Hierarchical Visual-Topological Memory (VTM)
  • VTM indexes visual memories by coarse room topology
  • Distinguishes in-room from remote-visible evidence
  • Retains successful approach cues across episodes
  • Re-localizes agent in accumulated scene structure
  • Retrieves target-relevant records from plausible rooms
  • Operates with fixed model parameters and no retraining

Entities

Sources