55 citations · 209 across the 37 of their papers we have counts for
8 papers · 1 filter
Replay-enhanced Continual Reinforcement Learning
Tiantian Zhang, Kevin Zehua Shen, Zichuan Lin +4
Replaying past experiences has proven to be a highly effective approach for averting catastrophic forgetting in supervised continual learning. However, some crucial factors are sti…
Master-slave Deep Architecture for Top-K Multi-armed Bandits with Non-linear Bandit Feedback and Diversity Constraints
Hanchi Huang, Li Shen, Deheng Ye +1
We propose a novel master-slave architecture to solve the top- combinatorial multi-armed bandits problem with non-linear bandit feedback and diversity constraints, which, to the…
RLTF: Reinforcement Learning from Unit Test Feedback
Jiate Liu, Yiqin Zhu, Kaiwen Xiao +4
The goal of program synthesis, or code generation, is to generate executable code based on given descriptions. Recently, there has been an increasing number of studies employing re…
Future-conditioned Unsupervised Pretraining for Decision Transformer
Zhihui Xie, Zichuan Lin, Deheng Ye +3
Recent research in offline reinforcement learning (RL) has demonstrated that return-conditioned supervised learning is a powerful paradigm for decision-making problems. While promi…
Deploying Offline Reinforcement Learning with Human Feedback
Ziniu Li, Ke Xu, Liu Liu +3
Reinforcement learning (RL) has shown promise for decision-making tasks in real-world applications. One practical framework involves training parameterized policy models from an of…
Revisiting Estimation Bias in Policy Gradients for Deep Reinforcement Learning
Haoxuan Pan, Deheng Ye, Xiaoming Duan +4
We revisit the estimation bias in policy gradients for the discounted episodic Markov decision process (MDP) from Deep Reinforcement Learning (DRL) perspective. The objective is fo…