Exploration and Regularization of the Latent Action Space in Recommendation
arXiv:2302.03431 · doi:10.1145/3543507.3583244
Abstract
In recommender systems, reinforcement learning solutions have effectively boosted recommendation performance because of their ability to capture long-term user-system interaction. However, the action space of the recommendation policy is a list of items, which could be extremely large with a dynamic candidate item pool. To overcome this challenge, we propose a hyper-actor and critic learning framework where the policy decomposes the item list generation process into a hyper-action inference step and an effect-action selection step. The first step maps the given state space into a vectorized hyper-action space, and the second step selects the item list based on the hyper-action. In order to regulate the discrepancy between the two action spaces, we design an alignment module along with a kernel mapping function for items to ensure inference accuracy and include a supervision module to stabilize the learning process. We build simulated environments on public datasets and empirically show that our framework is superior in recommendation compared to standard RL baselines.
Proceedings of the ACM Web Conference 2023 (WWW '23), May 1--5, 2023, Austin, TX, USA
References in corpus (7)
- Reinforcement Knowledge Graph Reasoning for Explainable Recommendation
- Towards Long-term Fairness in Recommendation
- KuaiRand: An Unbiased Sequential Recommendation Dataset with Randomly Exposed Videos
- Toward Pareto Efficient Fairness-Utility Trade-off inRecommendation through Reinforcement Learning
- Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
- AutoLossGen: Automatic Loss Function Generation for Recommender Systems
- Constrained Reinforcement Learning for Short Video Recommendation
Cited by in corpus (9)
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender Systems
- Generative Flow Network for Listwise Recommendation
- Efficient and Robust Regularized Federated Recommendation
- Modeling User Retention through Generative Flow Networks
- Value Function Decomposition in Markov Recommendation Process
- Empowering Denoising Sequential Recommendation with Large Language Model Embeddings
- LLM-Enhanced Reinforcement Learning for Long-Term User Satisfaction in Interactive Recommendation
- SPARK: Adaptive Low-Rank Knowledge Graph Modeling in Hybrid Geometric Spaces for Recommendation
- Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed