ARTFEED — Contemporary Art Intelligence

SheetSage-A2S Dataset Enhances Audio-to-Score Transcription for Popular Music

other · 2026-08-07

The SheetSage-A2S Dataset has been launched by researchers as a pioneering resource for audio-to-score (A2S) studies in popular music. This dataset comprises 61 hours of audio featuring **kern score encodings for 9,468 segments derived from 6,066 distinct songs. The research team enhanced current A2S methodologies through data augmentation and the use of MuQ, a pretrained model designed for music audio feature extraction, improving generalization and feature relevance. Their model recorded a symbol error rate (SER) of 4.98% on the Quartets collection in classical music, which is a notable improvement over the previous best of 15.3% SER. For the SheetSage-A2S dataset, the model achieved a SER of 20.92%, addressing the overlooked A2S application in popular music compared to classical music.

Key facts

  • SheetSage-A2S Dataset includes 61 hours of audio with **kern score encodings for 9,468 clips from 6,066 unique songs.
  • It is the first dataset to facilitate audio-to-score (A2S) research for popular music.
  • The proposed A2S model uses data augmentation and MuQ, a pretrained feature-extraction model for music audio.
  • The model achieves 4.98% symbol error rate (SER) on the Quartets collection for classical music.
  • The previous state-of-the-art achieved 15.3% SER on the same classical music collection.
  • On the SheetSage-A2S dataset for popular music, the model achieves 20.92% SER.
  • Existing A2S systems primarily focus on classical music, leaving popular music underexplored.
  • The paper is available on arXiv under the identifier 2608.06165.

Entities

Institutions

  • arXiv

Sources