3 papers
cs.IR2023
MDDL: A Framework for Reinforcement Learning-based Position Allocation in Multi-Channel Feed
Xiaowen Shi, Ze Wang, Yuanying Cai +6
Nowadays, the mainstream approach in position allocation system is to utilize a reinforcement learning model to allocate appropriate locations for items in various channels and the…
stat.ML2022
TD3 with Reverse KL Regularizer for Offline Reinforcement Learning from Mixed Datasets
Yuanying Cai, Chuheng Zhang, Li Zhao +6
We consider an offline reinforcement learning (RL) setting where the agent need to learn from a dataset collected by rolling out multiple behavior policies. There are two challenge…
cs.LG2020
Exploration by Maximizing Rényi Entropy for Reward-Free RL Framework
Chuheng Zhang, Yuanying Cai, Longbo Huang +1
Exploration is essential for reinforcement learning (RL). To face the challenges of exploration, we consider a reward-free RL framework that completely separates exploration from e…