2 papers
cs.LG2018
Policy Optimization with Model-based Explorations
Feiyang Pan, Qingpeng Cai, An-Xiang Zeng +5
Model-free reinforcement learning methods such as the Proximal Policy Optimization algorithm (PPO) have successfully applied in complex decision-making problems such as Atari games…
cs.LG2018
Speeding up the Metabolism in E-commerce by Reinforcement Mechanism Design
Hua-Lin He, Chun-Xiang Pan, Qing Da +1
In a large E-commerce platform, all the participants compete for impressions under the allocation mechanism of the platform. Existing methods mainly focus on the short-term return…