When the Label Ignores the Request: Auditing Policy-Selected Targets in Synthetic Conversational Music Recommendation
arXiv:2609.39696 · doi:10.1145/3842413.3842427
Abstract
Synthetic dialogues generated by LLM pipelines now serve as complete conversational-recommendation benchmarks: an LLM listener talks to an LLM recommender, and the track logged next in the conversation becomes the official label for each turn. These policy-selected labels make large-scale evaluation reproducible, but they are proxies for what the simulated user asked. We audit the one place where label and request are directly comparable: turns where the user asks for an exact song by name. In the RecSys Challenge 2026 TalkPlay benchmark, using visible dialogue and catalog metadata alone, we find that the official label contradicts the user's exact-song request in half of the audited development turns. This matters beyond one benchmark: naming the desired item is the dominant intent in real music search, where deployed systems avoid substituting an alternative for an exactly named item, on the premise that it costs satisfaction. A small training-time supplement closes most of the gap: adding catalog-resolved request-satisfying targets to a small fraction of training turns yields a 53.3% relative gain in nDCG@20 on the 43 conflict turns while leaving the official metric intact, verified against a matched control that detects the same requests but trains only on official labels.
6 pages, 2 tables. Camera-ready version (CC BY 4.0). RecSys Challenge 2026 Workshop at ACM RecSys 2026, Minneapolis, October 2, 2026. Code and audit artifacts: https://github.com/Sanjeev-S/recsys2026-request-audit
References in corpus (10)
- XGBoost: A Scalable Tree Boosting System
- Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches
- How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility
- Denoising Implicit Feedback for Recommendation
- Perspectives on Large Language Models for Relevance Judgment
- Large Language Models as Zero-Shot Conversational Recommenders
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
- Beyond NDCG: behavioral testing of recommender systems with RecList
- Limitations of Current Evaluation Practices for Conversational Recommender Systems and the Potential of User Simulation
- A Standardized Re-evaluation of Conversational Recommender Systems on the ReDial Dataset