Cross-Modal Bootstrapping Learns Implicit Music Style for Piano Arrangement
A recent paper published on arXiv (2608.03050) presents a novel cross-modal framework designed to learn implicit music styles from unprocessed audio and apply them to the generation of symbolic music, particularly for piano compositions. Drawing inspiration from BLIP-2, the model employs a Querying Transformer (Q-Former) to derive style representations from a pre-trained audio language model, subsequently conditioning a symbolic language model to create piano performances. This two-phase training approach incorporates contrastive learning to synchronize auditory style with symbolic representation, followed by generative modeling. The system produces piano arrangements based on both a lead sheet (content) and a reference audio piece (style), facilitating controllable and stylistically accurate outputs. While the experiments validate the method's effectiveness, specific results are not discussed in the abstract. The authors, researchers in the field, have submitted this work to arXiv, addressing the complexities of defining music style, which is often conveyed through text labels but remains implicit in tangible examples. The framework seeks to directly capture these implicit styles from audio, introducing a fresh approach to style-conditioned music generation.
Key facts
- Paper arXiv:2608.03050 introduces a cross-modal framework for learning music style from raw audio.
- The model is inspired by BLIP-2 and uses a Querying Transformer (Q-Former).
- Style representations are extracted from a pre-trained audio language model.
- The framework applies learned styles to symbolic music generation for piano arrangements.
- Two-stage training: contrastive learning and generative modeling.
- Piano performances are generated conditioned on a lead sheet and a reference audio example.
- The approach enables controllable and stylistically faithful arrangement.
- The paper is available on arXiv.
Entities
Institutions
- arXiv