ARTFEED — Contemporary Art Intelligence

IM-LEPP: A Hierarchical Energy-Based Model for Multimodal Cognition

ai-technology · 2026-08-15

A recent study published on arXiv (2608.12398) presents IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical energy-based model that enhances the single-modality LEPP framework by incorporating both vision and language. This model views cognition as latent states navigating through learned energy landscapes, paralleling concepts from statistical mechanics and thermodynamics. It features a hub-and-spoke architecture based on the controlled semantic cognition framework proposed by Lambon Ralph et al., where predictive-coding pathways for visual objects, scenes, and linguistic elements converge at a common amodal hub resembling the anterior temporal lobe. Importantly, predictions from each pathway are influenced by the hub state, maintaining unique pipeline information. The paper is classified as a cross-type announcement and can be accessed at https://arxiv.org/abs/2608.12398.

Key facts

  • IM-LEPP extends the single-modality LEPP model to integrate vision and language.
  • The model is hierarchical and energy-based, treating cognition as latent states flowing through energy landscapes.
  • It draws an analogy between generative neural networks and statistical mechanics.
  • The architecture uses a hub-and-spoke hierarchy based on the controlled semantic cognition framework of Lambon Ralph et al.
  • Predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub.
  • The hub is modeled on the anterior temporal lobe.
  • Each pipeline's prediction is conditioned by the current hub state, not overwritten.
  • The paper is available on arXiv with ID 2608.12398.

Entities

Institutions

  • arXiv

Sources