Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language Models
arXiv:2305.13112 · doi:10.18653/v1/2023.emnlp-main.621
Abstract
The recent success of large language models (LLMs) has shown great potential to develop more powerful conversational recommender systems (CRSs), which rely on natural language conversations to satisfy user needs. In this paper, we embark on an investigation into the utilization of ChatGPT for conversational recommendation, revealing the inadequacy of the existing evaluation protocol. It might over-emphasize the matching with the ground-truth items or utterances generated by human annotators, while neglecting the interactive nature of being a capable CRS. To overcome the limitation, we further propose an interactive Evaluation approach based on LLMs named iEvaLM that harnesses LLM-based user simulators. Our evaluation approach can simulate various interaction scenarios between users and systems. Through the experiments on two publicly available CRS datasets, we demonstrate notable improvements compared to the prevailing evaluation protocol. Furthermore, we emphasize the evaluation of explainability, and ChatGPT showcases persuasive explanation generation for its recommendations. Our study contributes to a deeper comprehension of the untapped potential of LLMs for CRSs and provides a more flexible and easy-to-use evaluation framework for future research endeavors. The codes and data are publicly available at https://github.com/RUCAIBox/iEvaLM-CRS.
Accepted by EMNLP 2023
References in corpus (19)
- Training language models to follow instructions with human feedback
- A Survey of Large Language Models
- A Survey on Accuracy-oriented Neural Recommendation: From Collaborative Filtering to Information-rich Recommendation
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Estimation-Action-Reflection: Towards Deep Interaction Between Conversational and Recommender Systems
- Uncovering ChatGPT's Capabilities in Recommender Systems
- Towards Unified Conversational Recommender Systems via Knowledge-Enhanced Prompt Learning
- Text and Code Embeddings by Contrastive Pre-Training
- Evaluating Mixed-initiative Conversational Search Systems via User Simulation
- Variational Reasoning over Incomplete Knowledge Graphs for Conversational Recommendation
- Is ChatGPT Equipped with Emotional Dialogue Capabilities?
- Improving Language Model Negotiation with Self-Play and In-Context Learning from AI Feedback
- Towards Explainable Conversational Recommender Systems
- UserSimCRS: A User Simulation Toolkit for Evaluating Conversational Recommender Systems
- Do LLMs Understand User Preferences? Evaluating LLMs On User Rating Prediction
- Improving Conversational Recommendation Systems via Counterfactual Data Simulation
- Measuring "Why" in Recommender Systems: a Comprehensive Survey on the Evaluation of Explainable Recommendation
- On the Tool Manipulation Capability of Open-source Large Language Models
- BARCOR: Towards A Unified Framework for Conversational Recommendation Systems
Cited by in corpus (8)
- Recommender Systems in the Era of Large Language Models (LLMs)
- When Large Language Models Meet Personalization: Perspectives of Challenges and Opportunities
- Large Language Models as Zero-Shot Conversational Recommenders
- Towards Empathetic Conversational Recommender Systems
- User Simulation for Evaluating Information Access Systems
- RecUserSim: A Realistic and Diverse User Simulator for Evaluating Conversational Recommender Systems
- CRS Arena: Crowdsourced Benchmarking of Conversational Recommender Systems
- Improving Conversational Recommendation with Contextual Adaptation of External Recommenders and LLM-based Reranking