collaborators

5 papers

cs.LG2025

Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR

Yuyang Zhang, Yang Hu, Bo Dai +1

Soft actor-critic (SAC) is a popular algorithm for max-entropy reinforcement learning. In practice, the energy-based policies in SAC are often approximated using simple policy clas…

cs.RO2025

Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations

Haitong Ma, Bo Dai, Zhaolin Ren +2

Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducin…

cs.LG2025

One-Step Flow Policy Mirror Descent

Tianyi Chen, Haitong Ma, Na Li +2

Diffusion policies have achieved great success in online reinforcement learning (RL) due to their strong expressive capacity. However, the inference of diffusion policy models reli…

cs.LG2025

Efficient Online Reinforcement Learning for Diffusion Policy

Haitong Ma, Tianyi Chen, Kai Wang +2

Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diff…

cs.MA2024

Scalable spectral representations for multi-agent reinforcement learning in network MDPs

Zhaolin Ren, Runyu Zhang, Bo Dai +1

Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global st…