1 paper
Wentao Shi, Xiangnan He, Yang Zhang +5
Planning for both immediate and long-term benefits becomes increasingly important in recommendation. Existing methods apply Reinforcement Learning (RL) to learn planning capacity b…