5 papers
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
Yuyang Zhang, Yang Hu, Bo Dai +1
Soft actor-critic (SAC) is a popular algorithm for max-entropy reinforcement learning. In practice, the energy-based policies in SAC are often approximated using simple policy clas…
Offline Imitation Learning upon Arbitrary Demonstrations by Pre-Training Dynamics Representations
Haitong Ma, Bo Dai, Zhaolin Ren +2
Limited data has become a major bottleneck in scaling up offline imitation learning (IL). In this paper, we propose enhancing IL performance under limited expert data by introducin…
One-Step Flow Policy Mirror Descent
Tianyi Chen, Haitong Ma, Na Li +2
Diffusion policies have achieved great success in online reinforcement learning (RL) due to their strong expressive capacity. However, the inference of diffusion policy models reli…
Efficient Online Reinforcement Learning for Diffusion Policy
Haitong Ma, Tianyi Chen, Kai Wang +2
Diffusion policies have achieved superior performance in imitation learning and offline reinforcement learning (RL) due to their rich expressiveness. However, the conventional diff…
Scalable spectral representations for multi-agent reinforcement learning in network MDPs
Zhaolin Ren, Runyu Zhang, Bo Dai +1
Network Markov Decision Processes (MDPs), a popular model for multi-agent control, pose a significant challenge to efficient learning due to the exponential growth of the global st…