18 citations · 25 across the 2 of their papers we have counts for
3 papers
cs.LG2022★ 7 cited
You May Not Need Ratio Clipping in PPO
Mingfei Sun, Vitaly Kurin, Guoqing Liu +4
Proximal Policy Optimization (PPO) methods learn a policy by iteratively performing multiple mini-batch optimization epochs of a surrogate objective with one set of sampled data. R…
cs.LG2021★ 18 cited
Return-Based Contrastive Representation Learning for Reinforcement Learning
Guoqing Liu, Chuheng Zhang, Li Zhao +5
Recently, various auxiliary tasks have been proposed to accelerate representation learning and improve sample efficiency in deep reinforcement learning (RL). However, existing auxi…
cs.AI2020
Suphx: Mastering Mahjong with Deep Reinforcement Learning
Junjie Li, Sotetsu Koyamada, Qiwei Ye +7
Artificial Intelligence (AI) has achieved great success in many domains, and game AI is widely regarded as its beachhead since the dawn of AI. In recent years, studies on game AI h…