3 papers
cs.LG2026
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning
Haitong Ma, Chenxiao Gao, Tianyi Chen +2
A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induces an overgreedy policy and fail…
cs.RO2025
Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
Haitong Ma, Bo Dai, Zhaolin Ren +2
Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducin…
cs.LG2024
Distributed Thompson sampling under constrained communication
Saba Zerefa, Zhaolin Ren, Haitong Ma +1
In Bayesian optimization, a black-box function is maximized via the use of a surrogate model. We apply distributed Thompson sampling, using a Gaussian process as a surrogate model,…