1 citations · 1 across the 3 of their papers we have counts for
3 papers
cs.LG2025
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
Hao Hu, Xinqi Wang, Simon Shaolei Du
We introduce a novel task of clustering trajectories from offline reinforcement learning (RL) datasets, where each cluster center represents the policy that generated its trajector…
cs.LG2025
CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
Ni Mu, Hao Hu, Xiao Hu +3
Preference-based reinforcement learning (PbRL) bypasses explicit reward engineering by inferring reward functions from human preference comparisons, enabling better alignment with…
cs.LG2023★ 1 cited
What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?
Rui Yang, Yong Lin, Xiaoteng Ma +3
Offline goal-conditioned RL (GCRL) offers a way to train general-purpose agents from fully offline datasets. In addition to being conservative within the dataset, the generalizatio…