BAMM Dataset Exposes AI Music Detection Gaps in Real Broadcasts
A new research study has been released on arXiv, presenting BAMM, which stands for Broadcast AI-Music Monitoring. This dataset comprises 40 hours of actual television footage featuring music created by both artificial intelligence and human artists. The study tackles the complex challenge of detecting AI-generated music in real-time broadcasts, as earlier techniques were reliant on synthetic datasets. Researchers evaluated CNN models using two training approaches—clean and broadcast—across three scenarios: Clean Foreground Music, Synthetic TV Broadcast, and Real TV Broadcast. While models performed well in clean conditions, accuracy significantly declined in synthetic environments, illustrating the difficulties in identifying AI music in regular media.
Key facts
- BAMM is a 40-hour dataset of real-world television recordings.
- Dataset includes AI-generated and human-made music.
- Study compares clean-trained and broadcast-trained CNN variants.
- Three scenarios: Clean Foreground Music (CFM), Synthetic TV Broadcast (STB), Real TV Broadcast (RTB).
- Both models achieve near-perfect performance on CFM.
- Performance degrades substantially under synthetic broadcast conditions.
- Broadcast-oriented training improves robustness but performance remains limited.
- Evaluation on RTB reveals further challenges.
- Paper available on arXiv with identifier 2608.07359.
- Concerns about transparency and fair compensation in broadcast media.
Entities
Institutions
- arXiv