dots.tts.edit: Precise Speech Editing with Continuous Autoregressive Model
A recent scholarly article presents dots.tts.edit, a speech editing tool that employs transcript-based structural editing instructions featuring XML-like tags to define typed actions and confine them to specific transcript segments or limits. This system eliminates the need for direct timestamp synchronization and offers a verifiable framework for compositional modifications. The editing tool is derived from the continuous autoregressive dots.tts foundational model. It includes four key controls for speech creation, addressing lexical content, emotional tone, pitch and speaking rate, as well as temporal phrasing via text and em. The article can be accessed on arXiv under the identifier 2608.02673.
Key facts
- The paper is titled 'dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model'.
- It is available on arXiv with identifier 2608.02673.
- The system uses a transcript-grounded structural edit instruction with XML-style tags.
- The interface specifies typed operations and localizes them to transcript spans or boundaries.
- It avoids explicit timestamp alignment.
- The editor is adapted from the continuous autoregressive dots.tts foundation model.
- Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing.
- The paper is categorized as a cross-type announcement.
Entities
Institutions
- arXiv