ARTFEED — Contemporary Art Intelligence

Spatial Memory Agent Enhances VLM Spatial Reasoning Without Parameter Updates

ai-technology · 2026-08-15

A new research paper on arXiv (2608.12743) introduces Spatial Memory Agent (SMA), a runtime framework that improves spatial reasoning in vision-language models (VLMs) without updating model parameters or relying on external spatial tools at inference time. The paper, titled "Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence," addresses a gap in existing approaches: post-training methods (supervised fine-tuning, reinforcement learning) and agentic paradigms that call external tools like depth estimation and 3D reconstruction. SMA instead enables a frozen VLM to self-evolve by converting verified spatial experiences into reusable, transferable lessons. In verifiable spatial environments, SMA queries the frozen model, validates outcomes, and stores successful procedures in a memory bank, which can be retrieved and applied to new tasks. This parameter-update-free self-evolution offers a complementary route to enhance spatial intelligence for embodied agents, robotic planning, and multimodal assistants. The paper is available at https://arxiv.org/abs/2608.12743.

Key facts

  • Paper arXiv:2608.12743 introduces Spatial Memory Agent (SMA)
  • SMA improves VLM spatial reasoning without parameter updates
  • SMA does not depend on external spatial tools at inference time
  • SMA uses experience-grounded runtime framework with reusable lessons
  • SMA queries frozen VLM in verifiable spatial environments
  • Existing methods include post-training and agentic tool use
  • SMA targets embodied agents, robotic planning, and multimodal assistants
  • Paper is available on arXiv

Entities

Institutions

  • arXiv

Sources