ARTFEED — Contemporary Art Intelligence

MPEcho: AI Framework for Controllable Cover Song Generation

ai-technology · 2026-07-30

A novel AI framework named MPEcho enhances the generation of cover songs by incorporating explicit phoneme-level conditioning and accurate temporal boundaries into the established SongEcho model. The advanced SongEcho relies on F0 sequences and voiced/unvoiced tags, resulting in elevated phoneme error rates due to its reliance on implicit linguistic data. Drawing inspiration from singing voice synthesis, MPEcho integrates a phoneme encoder and a length regulator, leading to a notable decrease in phoneme error rates. To facilitate this, the team created Phonsa, an automatic transcription model based on Whisper, which delivers high-precision phoneme-level annotations for singing voices, addressing the lack of quality audio-phoneme pairs. Experimental findings confirm Phonsa's alignment effectiveness. The research is available on arXiv with ID 2607.26698.

Key facts

  • MPEcho is a generative framework for controllable cover song generation.
  • It preserves melodic and linguistic content while recreating other musical components.
  • The state-of-the-art model SongEcho uses F0 sequences and V/UV tags.
  • SongEcho's implicit linguistic information leads to high phoneme error rate (PER).
  • MPEcho integrates a phoneme encoder and length regulator into SongEcho.
  • MPEcho provides explicit phoneme-level conditioning and precise temporal boundaries.
  • Phonsa is a Whisper-based automatic transcription model for phoneme annotations.
  • Phonsa overcomes scarcity of high-quality audio-phoneme pairs for singing voices.

Entities

Institutions

  • arXiv

Sources