ARTFEED — Contemporary Art Intelligence

PRISM-AH Framework for Video-Level Ambivalence and Hesitancy Recognition

other · 2026-07-29

A new system called PRISM-AH, which stands for Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition, has been developed to detect ambivalence and hesitancy (A/H) in videos. A/H refers to mixed emotional responses that can hinder or prevent people from making health-related changes. This emotional conflict can stem from different facial expressions, vocal inflections, language use, and body movements, varying by individual. PRISM-AH views A/H as a complex conflict that changes over time. It combines visual, audio, and textual data in short segments, utilizing a lightweight model to evaluate these interactions, forecast future moments, and spot signs of hesitation. The research can be found on arXiv with ID 2607.25961.

Key facts

  • PRISM-AH stands for Predictive Reasoning over Interacting Streams for Multimodal Ambivalence/Hesitancy Recognition.
  • Ambivalence and hesitancy (A/H) are conflicting affective states that precede delay or abandonment of health behavior change.
  • Recognition of A/H at the video level is difficult due to disagreement across facial, vocal, linguistic, and bodily modalities.
  • PRISM-AH uses frozen vision, audio, and text encoders aligned into short time windows.
  • A lightweight streaming model scores cross-modal dissonance and predicts each next window.
  • The model discovers behavior prototypes and is conditioned on participant metadata.
  • Dense window-level annotations supervise the model as an auxiliary objective.
  • The research is published on arXiv with ID 2607.25961.

Entities

Institutions

  • arXiv

Sources