most citedDiscriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations

9 citations · 11 across the 3 of their papers we have counts for

collaborators

5 papers

cs.LG20235 cited

PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Jianxiong Li, Xiao Hu, Haoran Xu +3

Offline-to-online reinforcement learning (RL), by combining the benefits of offline pretraining and online finetuning, promises enhanced sample efficiency and policy performance. H…

cs.LG20234 cited

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Haoran Xu, Li Jiang, Jianxiong Li +4

Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the devi…

cs.LG20231 cited

Mind the Gap: Offline Policy Optimization for Imperfect Rewards

Jianxiong Li, Xiao Hu, Haoran Xu +4

Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to des…

cs.LG20229 cited

Discriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations

Haoran Xu, Xianyuan Zhan, Honglei Yin +1

We study the problem of offline Imitation Learning (IL) where an agent aims to learn an optimal expert behavior policy without additional online environment interactions. Instead,…

cs.CV20221 cited

Adversarial Contrastive Learning via Asymmetric InfoNCE

Qiying Yu, Jieming Lou, Xianyuan Zhan +4

Contrastive learning (CL) has recently been applied to adversarial learning tasks. Such practice considers adversarial samples as additional positive views of an instance, and by m…