Ex-Omni-2D: Omni-Modal Dialogue with Native Visual Presence
The Ex-Omni-2D framework has been developed to allow omni-modal dialogue models to produce synchronized outputs that include text, customized speech, and video conditioned on references. This innovative system, detailed in an arXiv paper (2608.10720), overcomes the shortcomings of current omni-modal models that generate audio responses without visual representation. When presented with a multimodal query, along with a reference image and audio, Ex-Omni-2D generates a structured Visual Thought Plan (VTP) outlining scene, emotion, and motion, followed by text and native multi-codebook speech units. These units create a unified acoustic-temporal interface, which is decoded into speech and synchronized with video frames. This setup enables learning from diverse speech, dialogue, and avatar-video datasets without extensive supervision. A full-sequence Video Generator acts as the main instructor for effective training. The paper was recently submitted to arXiv.
Key facts
- Ex-Omni-2D is an omni-modal dialogue framework that generates text, speech, and video responses.
- It uses a Visual Thought Plan (VTP) to describe scene, emotion, and motion.
- The system employs native multi-codebook speech units as a shared acoustic-temporal interface.
- It aligns speech with video frames online.
- Training leverages heterogeneous speech, dialogue, and avatar-video data.
- A full-sequence Video Generator acts as the primary teacher.
- The paper is arXiv:2608.10720.
- The framework aims to provide native visual presence in omni-modal dialogue.
Entities
Institutions
- arXiv