5 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 5 cited
PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning
Jianxiong Li, Xiao Hu, Haoran Xu +3
Offline-to-online reinforcement learning (RL), by combining the benefits of offline pretraining and online finetuning, promises enhanced sample efficiency and policy performance. H…
cs.LG2023★ 1 cited
Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Jianxiong Li, Xiao Hu, Haoran Xu +4
Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to des…