MPEcho: AI Framework for Controllable Cover Song Generation
A novel AI framework named MPEcho enhances the generation of cover songs by incorporating explicit phoneme-level conditioning and accurate temporal boundaries into the established SongEcho model. The advanced SongEcho relies on F0 sequences and voiced/unvoiced tags, resulting in elevated phoneme error rates due to its reliance on implicit linguistic data. Drawing inspiration from singing voice synthesis, MPEcho integrates a phoneme encoder and a length regulator, leading to a notable decrease in phoneme error rates. To facilitate this, the team created Phonsa, an automatic transcription model based on Whisper, which delivers high-precision phoneme-level annotations for singing voices, addressing the lack of quality audio-phoneme pairs. Experimental findings confirm Phonsa's alignment effectiveness. The research is available on arXiv with ID 2607.26698.
Key facts
- MPEcho is a generative framework for controllable cover song generation.
- It preserves melodic and linguistic content while recreating other musical components.
- The state-of-the-art model SongEcho uses F0 sequences and V/UV tags.
- SongEcho's implicit linguistic information leads to high phoneme error rate (PER).
- MPEcho integrates a phoneme encoder and length regulator into SongEcho.
- MPEcho provides explicit phoneme-level conditioning and precise temporal boundaries.
- Phonsa is a Whisper-based automatic transcription model for phoneme annotations.
- Phonsa overcomes scarcity of high-quality audio-phoneme pairs for singing voices.
Entities
Institutions
- arXiv