C$^2$MOE: A New Framework for Incomplete Multimodal Emotion Recognition
A new framework called C$^2$MOE has been developed by researchers to tackle the issue of absent modalities in Multimodal Emotion Recognition in Conversations (MERC). This framework, outlined in a paper on arXiv (arXiv:2608.04013), integrates representation learning with missing modality imputation through an information-theoretic lens. It distinguishes multimodal knowledge into components of consistency and complementarity using interaction-aware experts. Consistency is enhanced by maximizing predictability across modalities, while complementarity is maintained to prevent biased reconstructions. This method seeks to bolster model resilience in practical situations where missing modalities arise from transmission issues or user actions. The paper notes that while current techniques focus on cross-modal consistency, they often overlook modality complementarity, resulting in biased outputs. C$^2$MOE aims to rectify this gap by explicitly addressing both elements, potentially advancing emotion recognition in conversational AI.
Key facts
- C$^2$MOE is a Consistency and Complementarity-guided Mixture of Experts framework.
- It targets incomplete multimodal emotion learning in Multimodal Emotion Recognition in Conversations (MERC).
- The framework unifies representation learning and missing modality imputation.
- It uses an information-theoretic framework.
- Multimodal knowledge is factorized into consistency and complementarity components.
- Consistency is captured by maximizing cross-modal predictability.
- Complementarity is preserved to avoid biased reconstructions.
- The paper is available on arXiv with identifier 2608.04013.
Entities
Institutions
- arXiv