activity
20192025
most citedSupervised Learning Achieves Human-Level Performance in MOBA Games: A Case Study of Honor of Kings

55 citations · 209 across the 37 of their papers we have counts for

collaborators
Showing 2023Show all

8 papers · 1 filter

cs.LG2023★ 3 cited

Replay-enhanced Continual Reinforcement Learning

Tiantian Zhang, Kevin Zehua Shen, Zichuan Lin +4

Replaying past experiences has proven to be a highly effective approach for averting catastrophic forgetting in supervised continual learning. However, some crucial factors are sti…

cs.LG2023

Master-slave Deep Architecture for Top-K Multi-armed Bandits with Non-linear Bandit Feedback and Diversity Constraints

Hanchi Huang, Li Shen, Deheng Ye +1

We propose a novel master-slave architecture to solve the top- combinatorial multi-armed bandits problem with non-linear bandit feedback and diversity constraints, which, to the…

cs.AI2023★ 1 cited

RLTF: Reinforcement Learning from Unit Test Feedback

Jiate Liu, Yiqin Zhu, Kaiwen Xiao +4

The goal of program synthesis, or code generation, is to generate executable code based on given descriptions. Recently, there has been an increasing number of studies employing re…

cs.LG2023★ 2 cited

Future-conditioned Unsupervised Pretraining for Decision Transformer

Zhihui Xie, Zichuan Lin, Deheng Ye +3

Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…

cs.LG2023★ 2 cited

Deploying Offline Reinforcement Learning with Human Feedback

Ziniu Li, Ke Xu, Liu Liu +3

Reinforcement learning (RL) has shown promise for decision-making tasks in real-world applications. One practical framework involves training parameterized policy models from an of…

cs.LG2023

Revisiting Estimation Bias in Policy Gradients for Deep Reinforcement Learning

Haoxuan Pan, Deheng Ye, Xiaoming Duan +4

We revisit the estimation bias in policy gradients for the discounted episodic Markov decision process (MDP) from Deep Reinforcement Learning (DRL) perspective. The objective is fo…