IM-LEPP: A Hierarchical Energy-Based Model for Multimodal Cognition
A recent study published on arXiv (2608.12398) presents IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical energy-based model that enhances the single-modality LEPP framework by incorporating both vision and language. This model views cognition as latent states navigating through learned energy landscapes, paralleling concepts from statistical mechanics and thermodynamics. It features a hub-and-spoke architecture based on the controlled semantic cognition framework proposed by Lambon Ralph et al., where predictive-coding pathways for visual objects, scenes, and linguistic elements converge at a common amodal hub resembling the anterior temporal lobe. Importantly, predictions from each pathway are influenced by the hub state, maintaining unique pipeline information. The paper is classified as a cross-type announcement and can be accessed at https://arxiv.org/abs/2608.12398.
Key facts
- IM-LEPP extends the single-modality LEPP model to integrate vision and language.
- The model is hierarchical and energy-based, treating cognition as latent states flowing through energy landscapes.
- It draws an analogy between generative neural networks and statistical mechanics.
- The architecture uses a hub-and-spoke hierarchy based on the controlled semantic cognition framework of Lambon Ralph et al.
- Predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub.
- The hub is modeled on the anterior temporal lobe.
- Each pipeline's prediction is conditioned by the current hub state, not overwritten.
- The paper is available on arXiv with ID 2608.12398.
Entities
Institutions
- arXiv