Can Offline Metrics Measure Explanation Goals? A Comparative Survey Analysis of Offline Explanation Metrics in Recommender Systems
arXiv:2310.14379 · doi:10.1145/3779420
Abstract
In Recommender System (RS), explanations help users understand why items are recommended and can enhance a system's transparency, persuasiveness, engagement, and trust, which are known as explanation goals. However, evaluating the effectiveness of explanation algorithms offline remains challenging because explanation goals are inherently subjective. We initially conducted a rapid literature review, which revealed that algorithms are often assessed using anecdotal evidence (offering convincing examples) or using metrics that do not align with human perception. From these results, we investigated whether the selection of item attributes and interacted items affects explanation goals in explanations that generate a path connecting interacted and recommended items based on shared attributes (such as genres). We used metrics that measure the diversity and popularity of attributes and the recency of item interactions to evaluate explanations from three state-of-the-art agnostic algorithms across six recommendation systems. We then performed an online user study to compare user perceptions of explanation goals and offline metrics. Our findings indicate that engagement is sensitive to users' perceptions of diversity in explanations, whereas transparency, trust, and persuasiveness are influenced by perceptions of both popularity and diversity. However, offline metrics require refinement to more closely align with explanation goals and user understanding.
References in corpus (13)
- KGAT: Knowledge Graph Attention Network for Recommendation
- Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making
- Unifying Knowledge Graph Learning and Recommendation: Towards a Better Understanding of User Preferences
- Disentangled Graph Collaborative Filtering
- Learning Intents behind Interactions with Knowledge Graph for Recommendation
- From Anecdotal Evidence to Quantitative Evaluation Methods: A Systematic Review on Evaluating Explainable AI
- Reinforcement Knowledge Graph Reasoning for Explainable Recommendation
- A systematic review and taxonomy of explanations in decision support and recommender systems
- Embarrassingly Shallow Autoencoders for Sparse Data
- Post Processing Recommender Systems with Knowledge Graphs for Recency, Popularity, and Diversity of Explanations
- Finding Paths for Explainable MOOC Recommendation: A Learner Perspective
- Towards Next-Generation LLM-based Recommender Systems: A Survey and Beyond
- Sustainable transparency in Recommender Systems: Bayesian Ranking of Images for Explainability