2 papers
cs.LG2024
Transferable Reward Learning by Dynamics-Agnostic Discriminator Ensemble
Fan-Ming Luo, Xingchen Cao, Rong-Jun Qin +1
Recovering reward function from expert demonstrations is a fundamental problem in reinforcement learning. The recovered reward function captures the motivation of the expert. Agent…
cs.LG2024
Efficient Recurrent Off-Policy RL Requires a Context-Encoder-Specific Learning Rate
Fan-Ming Luo, Zuolin Tu, Zefang Huang +1
Real-world decision-making tasks are usually partially observable Markov decision processes (POMDPs), where the state is not fully observable. Recent progress has demonstrated that…