26 citations · 27 across the 5 of their papers we have counts for
5 papers
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
Fan-Ming Luo, Zuolin Tu, Zefang Huang +1
Real-world decision-making tasks are usually partially observable Markov decision processes (POMDPs), where the state is not fully observable. Recent progress has demonstrated that…
Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement Learning
Fan-Ming Luo, Tian Xu, Xingchen Cao +1
Learning a precise dynamics model can be crucial for offline reinforcement learning, which, unfortunately, has been found to be quite challenging. Dynamics models that are learned…
Unified Policy Optimization for Continuous-action Reinforcement Learning in Non-stationary Tasks and Games
Rong-Jun Qin, Fan-Ming Luo, Hong Qian +1
This paper addresses policy learning in non-stationary environments and games with continuous actions. Rather than the classical reward maximization mechanism, inspired by the idea…
A Survey on Model-based Reinforcement Learning
Fan-Ming Luo, Tian Xu, Hang Lai +3
Reinforcement learning (RL) solves sequential decision-making problems via a trial-and-error process interacting with the environment. While RL achieves outstanding success in play…
Transferable Reward Learning by Dynamics-Agnostic Discriminator Ensemble
Fan-Ming Luo, Xingchen Cao, Rong-Jun Qin +1
Recovering reward function from expert demonstrations is a fundamental problem in reinforcement learning. The recovered reward function captures the motivation of the expert. Agent…