ARTFEED — Contemporary Art Intelligence

Rationale-Guided Learning Enhances Multimodal Emotion Recognition

ai-technology · 2026-08-13

A novel approach known as rationale-guided learning (RGL) has been introduced to enhance multimodal emotion recognition in conversation (MERC). This method differs from conventional techniques that link multimodal signals directly to emotion labels by integrating causal reasoning rooted in dual-process theory. RGL breaks down emotional reasoning into three components: Intuitive (System 1), Contextual (System 2), and Integrative. An offline multimodal large language model (MLLM) creates structured rationales, which are then stored as memories to facilitate model training, ensuring that internal representations mirror human reasoning. During inference, the final model functions without the MLLM's overhead. The research paper can be found on arXiv under ID 2608.10448.

Key facts

  • The framework is called rationale-guided learning (RGL).
  • It is designed for multimodal emotion recognition in conversation (MERC).
  • RGL is based on dual-process theory.
  • Emotional reasoning is decomposed into Intuitive, Contextual, and Integrative facets.
  • An offline MLLM generates structured rationales.
  • Rationales are encoded as memories to guide model training.
  • The final model requires no MLLM overhead at inference time.
  • The paper is available on arXiv with ID 2608.10448.

Entities

Institutions

  • arXiv

Sources