Partially Observable Markov Decision Process for Recommender Systems
arXiv:1608.07793
Abstract
We report the "Recurrent Deterioration" (RD) phenomenon observed in online recommender systems. The RD phenomenon is reflected by the trend of performance degradation when the recommendation model is always trained based on users' feedbacks of the previous recommendations. There are several reasons for the recommender systems to encounter the RD phenomenon, including the lack of negative training data and the evolution of users' interests, etc. Motivated to tackle the problems causing the RD phenomenon, we propose the POMDP-Rec framework, which is a neural-optimized Partially Observable Markov Decision Process algorithm for recommender systems. We show that the POMDP-Rec framework effectively uses the accumulated historical data from real-world recommender systems and automatically achieves comparable results with those models fine-tuned exhaustively by domain exports on public datasets.
References in corpus (6)
- Continuous control with deep reinforcement learning
- Playing Atari with Deep Reinforcement Learning
- A Contextual-Bandit Approach to Personalized News Article Recommendation
- Deep Recurrent Q-Learning for Partially Observable MDPs
- Value-Function Approximations for Partially Observable Markov Decision Processes
- Exploration vs. Exploitation in the Information Filtering Problem
Cited by in corpus (5)
- Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
- Reinforcement Learning to Optimize Long-term User Engagement in Recommender Systems
- Recovering Markov Models from Closed-Loop Data
- Offline Meta-level Model-based Reinforcement Learning Approach for Cold-Start Recommendation
- Novel Approaches to Accelerating the Convergence Rate of Markov Decision Process for Search Result Diversification