19 citations · 29 across the 4 of their papers we have counts for
11 papers
Constrained Reinforcement Learning for Short Video Recommendation
Qingpeng Cai, Ruohan Zhan, Chi Zhang +5
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users provide complex and…
Softmax Deep Double Deterministic Policy Gradients
Ling Pan, Qingpeng Cai, Longbo Huang
A widely-used actor-critic reinforcement learning algorithm for continuous control, Deep Deterministic Policy Gradients (DDPG), suffers from the overestimation problem, which can n…
Generator and Critic: A Deep Reinforcement Learning Approach for Slate Re-ranking in E-commerce
Jianxiong Wei, Anxiang Zeng, Yueqiu Wu +3
The slate re-ranking problem considers the mutual influences between items to improve user satisfaction in e-commerce, compared with the point-wise ranking. Previous works either d…
Multi-Path Policy Optimization
Ling Pan, Qingpeng Cai, Longbo Huang
Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, partic…
Deterministic Value-Policy Gradients
Qingpeng Cai, Ling Pan, Pingzhong Tang
Reinforcement learning algorithms such as the deep deterministic policy gradient algorithm (DDPG) has been widely used in continuous control tasks. However, the model-free DDPG alg…
Reinforcement Learning Driven Heuristic Optimization
Qingpeng Cai, Will Hang, Azalia Mirhoseini +3
Heuristic algorithms such as simulated annealing, Concorde, and METIS are effective and widely used approaches to find solutions to combinatorial optimization problems. However, th…