48 citations · 64 across the 7 of their papers we have counts for
6 papers · 1 filter
JourneyFormer: Encoding Airbnb Guest Journey with Sequence Modeling
Daochen Zha, Chun How Tan, Xin Liu +9
Sequence modeling has become increasingly popular in recommendation and ranking algorithms, owing to its capacity to model users' historical behaviors and infer user intentions. De…
STRAPPER: Preference-based Reinforcement Learning via Self-training Augmentation and Peer Regularization
Yachen Kang, Li He, Jinxin Liu +2
Preference-based reinforcement learning (PbRL) promises to learn a complex reward function with binary human preference. However, such human-in-the-loop formulation requires consid…
Beyond Reward: Offline Preference-guided Policy Optimization
Yachen Kang, Diyuan Shi, Jinxin Liu +2
This study focuses on the topic of offline preference-based reinforcement learning (PbRL), a variant of conventional reinforcement learning that dispenses with the need for online…
CEIL: Generalized Contextual Imitation Learning
Jinxin Liu, Li He, Yachen Kang +3
In this paper, we present \textbf{C}ont\textbf{E}xtual \textbf{I}mitation \textbf{L}earning~(CEIL), a general and broadly applicable algorithm for imitation learning (IL). Inspired…
CLUE: Calibrated Latent Guidance for Offline Reinforcement Learning
Jinxin Liu, Lipeng Zu, Li He +1
Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected and labeled datasets, which eliminates the time-consuming data collection in online RL. Howev…
OER: Offline Experience Replay for Continual Offline Reinforcement Learning
Sibo Gai, Donglin Wang, Li He
The capability of continuously learning new skills via a sequence of pre-collected offline datasets is desired for an agent. However, consecutively learning a sequence of offline t…