ARTFEED — Contemporary Art Intelligence

X2C Benchmark: 100,000 Pairs for Humanoid Facial Expression Imitation

ai-technology · 2026-08-13

A new benchmark dataset named X2C has been unveiled by researchers, comprising 100,000 pairs of (image, control value) that depict humanoid facial expressions, each annotated with 30 continuous control parameters. This development aims to bridge the divide between biological facial dynamics and mechanical control, which has been a barrier in correlating visual cues with actuation signals. The X2CNet framework, detailed in a paper available on arXiv (arXiv:2505.11146), operates in two stages to separate visual motion features from mechanical control regression. The study emphasizes progress in the visual synthesis of talking heads while acknowledging the difficulties in achieving accurate actuation. X2CNet's design facilitates the relationship between visual features and control parameters, benefiting fields such as computer vision, robotics, and human-computer interaction.

Key facts

  • X2C is a benchmark dataset with 100,000 (image, control value) pairs.
  • The dataset includes 30 continuous control parameters for nuanced expressions.
  • X2CNet is a two-stage deep learning framework for facial expression imitation.
  • The paper is available on arXiv with ID 2505.11146.
  • The announcement type is 'replace-cross'.
  • The work addresses the domain gap between biological facial dynamics and mechanical control spaces.
  • The benchmark aims to establish a high-fidelity standard for facial expression transfer.
  • The lack of large-scale paired data has been a primary obstacle.

Entities

Institutions

  • arXiv

Sources