5 papers · 1 filter
AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra +1
Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but e…
What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jiménez +2
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixe…
Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study
Simon Hachmeier, R. Oguz Araz, Dmitry Bogdanov +2
Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by pr…
Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
Pedro Ramoneda, Pablo Alonso-Jiménez, Sergio Oramas +2
Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descript…
Discogs-VI: A Musical Version Identification Dataset Based on Public Editorial Metadata
R. Oguz Araz, Xavier Serra, Dmitry Bogdanov
Current version identification (VI) datasets often lack sufficient size and musical diversity to train robust neural networks (NNs). Additionally, their non-representative clique s…