ARTFEED — Contemporary Art Intelligence

CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

ai-technology · 2026-08-18

A recent study presents CETalk, a novel framework designed for the creation of 3D talking heads with the ability to control emotions continuously. Published on arXiv under the identifier 2608.15110, this research tackles the shortcomings of current techniques that depend on fixed emotion categories, which do not adequately reflect the fluid nature of emotions. CETalk employs continuous Valence-Arousal (VA) models for precise emotional regulation. The framework forecasts a series of FLAME parameters through three main components: a Dynamic Emotion Modulation Module that adjusts emotional intensity based on audio cues, a Multi-Scale Temporal Modeling system that separates high-frequency speech movements from low-frequency emotional changes, and a Dynamic Fusion Mechanism that combines these diverse features. The study emphasizes the challenge posed by the mismatch in temporal frequency between audio articulation and emotional expression, making it significant for digital human animation and affective computing.

Key facts

  • CETalk is a framework for audio-driven 3D talking head generation with continuous emotion control.
  • It uses continuous Valence-Arousal (VA) representations instead of discrete emotion categories.
  • The framework predicts FLAME parameters for facial animation.
  • It includes a Dynamic Emotion Modulation Module, a Multi-Scale Temporal Modeling mechanism, and a Dynamic Fusion Mechanism.
  • The paper addresses the temporal frequency mismatch between audio articulation and emotional expression.
  • The research is published on arXiv with identifier 2608.15110.
  • The paper is categorized as a cross-type announcement.
  • The abstract is available at https://arxiv.org/abs/2608.15110.

Entities

Institutions

  • arXiv

Sources