ARTFEED — Contemporary Art Intelligence

Action Post-training Reduces VLM Depth Decodability: A Study of Molmo2-ER and MolmoAct2-LIBERO

ai-technology · 2026-08-17

A recent paper on arXiv (2608.08904) explores the influence of action post-training on the spatial comprehension of vision-language models (VLMs) when developing a vision-language-action model (VLA). The study examines depth perception, a fundamental aspect of spatiogeometric understanding, across each decoder layer of a matched open-source VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. Findings reveal that the VLA exhibits poorer depth decoding at all layers, a consistent deficit termed the 'floor.' Furthermore, this decline is inconsistent; while the base VLM's depth decodability enhances in its later layers, the VLA experiences a drop, referred to as the 'cliff.' The researchers identify late-layer MLP interference as the cause of this cliff, noting that removing late-layer MLP writes significantly improves decodability, unlike similar interventions in the base VLM. This study sheds light on the trade-offs associated with action post-training and its effects on perceptual skills.

Key facts

  • Paper arXiv:2608.08904v2, announce type replace-cross.
  • Probes depth perception from every decoder layer of Molmo2-ER (base VLM) and MolmoAct2-LIBERO (VLA).
  • VLA decodes depth worse at every layer, termed the 'floor'.
  • Base VLM improves depth decodability in final layers, VLA collapses, termed the 'cliff'.
  • Cliff causally localized to late-layer MLP interference.
  • Ablating late-layer MLP writes recovers majority of terminal decodability cliff.
  • Matched attention ablations and same intervention in base VLM produce no comparable recovery.
  • Module-level decomposition explains the dissociation.

Entities

Institutions

  • arXiv

Sources