PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
A recent paper published on arXiv (2608.05218) presents Phoneme-Driven Gaussian Splatting (PD-GS), a technique designed to enhance lip synchronization in audio-based talking-head rendering. This method builds upon 3D Gaussian Splatting (3DGS) by incorporating time-aligned phoneme tokens sourced from an automatic speech recognition (ASR) and forced-alignment framework. Central to this approach is the Linguistic Fusion Module (LFM), which dynamically merges these phoneme tokens to mitigate the 'leaky mouth' issue, where mouth movements become overly smoothed, breaching articulatory limits such as bilabial closures. The authors contend that existing techniques derive discrete articulatory actions from continuous acoustic signals, leading to biased predictions. PD-GS seeks to rectify this by offering precise phoneme-level direction.
Key facts
- Paper ID: arXiv:2608.05218
- Announcement type: new
- Proposes Phoneme-Driven Gaussian Splatting (PD-GS)
- Augments 3DGS with time-aligned phoneme tokens
- Uses ASR and forced-alignment pipeline
- Core component: Linguistic Fusion Module (LFM)
- Addresses 'leaky mouth' artifact in talking-head rendering
- Published on arXiv
Entities
Institutions
- arXiv