9 papers · 1 filter
GeMPO: Generalized Measure Matching for Online Diffusion Reinforcement Learning
Haitong Ma, Chenxiao Gao, Tianyi Chen +2
A commonly used family of RL algorithms for diffusion policies conducts softmax reweighting over samples from the behavior policy, which often induces an overgreedy policy and fail…
Spectral Ghost in Representation Learning: from Component Analysis to Self-Supervised Learning
Bo Dai, Na Li, Dale Schuurmans
Self-supervised learning (SSL) has improved empirical performance by unleashing the power of unlabeled data for practical applications. Specifically, SSL extracts the representatio…
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
Yuyang Zhang, Yang Hu, Bo Dai +1
Soft actor-critic (SAC) is a popular algorithm for max-entropy reinforcement learning. In practice, the energy-based policies in SAC are often approximated using simple policy clas…
One-Step Flow Policy Mirror Descent
Tianyi Chen, Haitong Ma, Na Li +2
Diffusion policies have achieved great success in online reinforcement learning (RL) due to their strong expressive capacity. However, the inference of diffusion policy models reli…
Efficient Online Reinforcement Learning for Diffusion Policy
Haitong Ma, Tianyi Chen, Kai Wang +2
Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diff…
Primal-Dual Spectral Representation for Off-policy Evaluation
Yang Hu, Tianyi Chen, Na Li +2
Off-policy evaluation (OPE) is one of the most fundamental problems in reinforcement learning (RL) to estimate the expected long-term payoff of a given target policy with only expe…