conversational recommendation 1human vs. model assessment 1large language models 1music recommendation 1response evaluation 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.IR2026
LLM-as-a-Judge for Evaluating System Responses in Conversational Music Recommendation
Seungheon Doh, Bruno Sguerra, Sergio Oramas +2
The paper investigates the reliability of using large language models as judges to evaluate the quality of responses generated by conversational music recommendation systems, compa…
cs.SD2025
Benchmarking Music Autotagging with MGPHot Expert Annotations vs. Generic Tag Datasets
Pedro Ramoneda, Pablo Alonso-Jiménez, Sergio Oramas +2
Music autotagging aims to automatically assign descriptive tags, such as genre, mood, or instrumentation, to audio recordings. Due to its challenges, diversity of semantic descript…