DialectS2S: New End-to-End Speech Dialogue Model for Low-Resource Chinese Dialects
Researchers have introduced DialectS2S, an end-to-end speech dialogue model designed specifically for low-resource Chinese dialects. The model addresses the scarcity of dialect speech data and the semantic inconsistency that arises during dialect adaptation. A scalable synthesis pipeline constructs training data, and a two-stage post-training strategy with self-aligned speech supervision aligns the semantic content of speech supervision with the model's evolved representations, improving speech stability and naturalness. The work is detailed in a paper on arXiv (2608.08067).
Key facts
- DialectS2S is an end-to-end speech dialogue model for Chinese dialects.
- It addresses low-resource dialect scenarios due to scarce dialect speech data.
- The model uses a scalable dialect speech dialogue synthesis pipeline for data construction.
- A two-stage post-training strategy with self-aligned speech supervision is introduced.
- The strategy aligns semantic content of speech supervision with evolved semantic representations.
- It aims to improve speech stability and naturalness in dialect generation.
- The paper is available on arXiv with ID 2608.08067.
- The announcement type is replace-cross.
Entities
Institutions
- arXiv