activity
20222024
most citedDiscriminator-Weighted Offline Imitation Learning from Suboptimal Demonstrations

9 citations · 20 across the 8 of their papers we have counts for

collaborators

8 papers

eess.IV2024

NTIRE 2024 Challenge on Short-form UGC Video Quality Assessment: Methods and Results

Xin Li, Kun Yuan, Yajing Pei +65

This paper reviews the NTIRE 2024 Challenge on Shortform UGC Video Quality Assessment (S-UGC VQA), where various excellent solutions are submitted and evaluated on the collected da…

cs.CL2023

Narrowing the Gap between Zero- and Few-shot Machine Translation by Matching Styles

Weiting Tan, Haoran Xu, Lingfeng Shen +5

Large language models trained primarily in a monolingual setting have demonstrated their ability to generalize to machine translation using zero- and few-shot examples with in-cont…

cs.LG20235 cited

PROTO: Iterative Policy Regularized Offline-to-Online Reinforcement Learning

Jianxiong Li, Xiao Hu, Haoran Xu +3

Offline-to-online reinforcement learning (RL), by combining the benefits of offline pretraining and online finetuning, promises enhanced sample efficiency and policy performance. H…

cs.LG20234 cited

Offline RL with No OOD Actions: In-Sample Learning via Implicit Value Regularization

Haoran Xu, Li Jiang, Jianxiong Li +4

Most offline reinforcement learning (RL) methods suffer from the trade-off between improving the policy to surpass the behavior policy and constraining the policy to limit the devi…

cs.CL2023

Language-Aware Multilingual Machine Translation with Self-Supervised Learning

Haoran Xu, Jean Maillard, Vedanuj Goswami

Multilingual machine translation (MMT) benefits from cross-lingual transfer but is a challenging multitask optimization problem. This is partly because there is no clear framework…

cs.LG20231 cited

Mind the Gap: Offline Policy Optimization for Imperfect Rewards

Jianxiong Li, Xiao Hu, Haoran Xu +4

Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to des…