Face-to-Speech Synthesis from Static Images via Style Diffusion
Researchers have introduced a Face-to-Speech (F2S) system capable of producing a realistic voice from a still facial image, eliminating the necessity for reference audio. A compact Face Adapter harmonizes facial recognition attributes with the style space of a static StyleTTS 2 model, utilizing soft-tuning on the upper layers of the face encoder. Tested on LRS3, an English TED-talk audiovisual dataset, the generated speech received UTMOS scores ranging from 3.7 to 4.0, either matching or surpassing the ground truth score of 3.61. Face-to-voice retrieval rates are above chance, and the synthesized voices align well with the intended speakers. Additionally, the English-trained adapter can fluently generate Spanish speech without any retraining, addressing zero-shot TTS challenges for historical figures or video game characters based solely on visual input.
Key facts
- Face-to-Speech framework predicts voice from static facial image
- Uses lightweight Face Adapter and soft-tuning of face encoder's upper blocks
- Aligns face-recognition features with style space of frozen StyleTTS 2 model
- Evaluated on LRS3 corpus of English TED-talk videos
- UTMOS scores 3.7-4.0, exceeding ground truth's 3.61
- Face-to-voice retrieval above chance
- Generated voice consistent with target speaker
- English-trained adapter produces fluent Spanish without retraining
Entities
—