9 citations · 11 across the 3 of their papers we have counts for
5 papers
PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning
Jianxiong Li, Xiao Hu, Haoran Xu +3
Offline-to-online reinforcement learning (RL), by combining the benefits of offline pretraining and online finetuning, promises enhanced sample efficiency and policy performance. H…
Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization
Haoran Xu, Li Jiang, Jianxiong Li +4
Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the devi…
Mind the Gap: Offline Policy Optimization for Imperfect Rewards
Jianxiong Li, Xiao Hu, Haoran Xu +4
Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to des…
Discriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations
Haoran Xu, Xianyuan Zhan, Honglei Yin +1
We study the problem of offline Imitation Learning (IL) where an agent aims to learn an optimal expert behavior policy without additional online environment interactions. Instead,…
Adversarial Contrastive Learning via Asymmetric InfoNCE
Qiying Yu, Jieming Lou, Xianyuan Zhan +4
Contrastive learning (CL) has recently been applied to adversarial learning tasks. Such practice considers adversarial samples as additional positive views of an instance, and by m…