PRISM-AH Framework for Video-Level Ambivalence and Hesitancy Recognition
A new system called PRISM-AH, which stands for Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition, has been developed to detect ambivalence and hesitancy (A/H) in videos. A/H refers to mixed emotional responses that can hinder or prevent people from making health-related changes. This emotional conflict can stem from different facial expressions, vocal inflections, language use, and body movements, varying by individual. PRISM-AH views A/H as a complex conflict that changes over time. It combines visual, audio, and textual data in short segments, utilizing a lightweight model to evaluate these interactions, forecast future moments, and spot signs of hesitation. The research can be found on arXiv with ID 2607.25961.
Key facts
- PRISM-AH stands for Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition.
- Ambivalence and hesitancy (A/H) are conflicting affective states that precede delay or abandonment of health behavior change.
- Recognition of A/H at the video level is difficult due to disagreement across facial, vocal, linguistic, and bodily modalities.
- PRISM-AH uses frozen vision, audio, and text encoders aligned into short time windows.
- A lightweight streaming model scores cross-modal dissonance and predicts each next window.
- The model discovers behavior prototypes and is conditioned on participant metadata.
- Dense window-level annotations supervise the model as an auxiliary objective.
- The research is published on arXiv with ID 2607.25961.
Entities
Institutions
- arXiv