Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
A recent preprint on arXiv (2607.29363) addresses the difficulties of achieving a balance between sequence length, representational capacity, and long-term stability in autoregressive (AR) audio and speech generation. The researchers suggest a co-design approach that integrates a low-frame-rate, high-dimensional, high-bandwidth continuous representation with a streaming generation framework, aiming for reliable high-fidelity reconstruction, enhanced single-token predictability, and improved long-horizon stability. This study breaks down the objective into two interconnected challenges: identifying the geometric and statistical characteristics of high-dimensional continuous tokens and creating a generation framework that can utilize them effectively. The paper is classified as a cross-type announcement and is accessible on arXiv.
Key facts
- Paper ID: arXiv:2607.29363
- Announcement type: cross
- Focus: autoregressive speech and audio generation
- Proposes low-frame-rate, high-dimensional continuous tokens
- Aims to improve reconstruction fidelity and generation quality
- Addresses distribution drift and AR error accumulation
- Available on arXiv
- Published in 2025 (implied by arXiv ID)
Entities
Institutions
- arXiv