8 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.SD2024
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
Ilaria Manco, Justin Salamon, Oriol Nieto
Audio-text contrastive models have become a powerful approach in music representation learning. Despite their empirical success, however, little is known about the influence of key…
cs.SD2024★ 8 cited
Foundation Models for Music: A Survey
Yinghao Ma, Anders Øland, Anton Ragni +39
In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This com…
cs.SD2023★ 7 cited
The Song Describer Dataset: a Corpus of Audio Captions for Music-and-Language Evaluation
Ilaria Manco, Benno Weck, SeungHeon Doh +10
We introduce the Song Describer dataset (SDD), a new crowdsourced corpus of high-quality audio-caption pairs, designed for the evaluation of music-and-language models. The dataset…