Spatial Memory Agent Enhances VLM Spatial Reasoning Without Parameter Updates
A new research paper on arXiv (2608.12743) introduces Spatial Memory Agent (SMA), a runtime framework that improves spatial reasoning in vision-language models (VLMs) without updating model parameters or relying on external spatial tools at inference time. The paper, titled "Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence," addresses a gap in existing approaches: post-training methods (supervised fine-tuning, reinforcement learning) and agentic paradigms that call external tools like depth estimation and 3D reconstruction. SMA instead enables a frozen VLM to self-evolve by converting verified spatial experiences into reusable, transferable lessons. In verifiable spatial environments, SMA queries the frozen model, validates outcomes, and stores successful procedures in a memory bank, which can be retrieved and applied to new tasks. This parameter-update-free self-evolution offers a complementary route to enhance spatial intelligence for embodied agents, robotic planning, and multimodal assistants. The paper is available at https://arxiv.org/abs/2608.12743.
Key facts
- Paper arXiv:2608.12743 introduces Spatial Memory Agent (SMA)
- SMA improves VLM spatial reasoning without parameter updates
- SMA does not depend on external spatial tools at inference time
- SMA uses experience-grounded runtime framework with reusable lessons
- SMA queries frozen VLM in verifiable spatial environments
- Existing methods include post-training and agentic tool use
- SMA targets embodied agents, robotic planning, and multimodal assistants
- Paper is available on arXiv
Entities
Institutions
- arXiv