4 papers · 1 filter
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
Yuyang Zhang, Yang Hu, Bo Dai +1
Soft actor-critic (SAC) is a popular algorithm for max-entropy reinforcement learning. In practice, the energy-based policies in SAC are often approximated using simple policy clas…
A Model-Based Approach to Imitation Learning through Multi-Step Predictions
Haldun Balim, Yang Hu, Yuyang Zhang +1
Imitation learning is a widely used approach for training agents to replicate expert behavior in complex decision-making tasks. However, existing methods often struggle with compou…
Primal-Dual Spectral Representation for Off-policy Evaluation
Yang Hu, Tianyi Chen, Na Li +2
Off-policy evaluation (OPE) is one of the most fundamental problems in reinforcement learning (RL) to estimate the expected long-term payoff of a given target policy with only expe…
Efficient Duple Perturbation Robustness in Low-rank MDPs
Yang Hu, Haitong Ma, Bo Dai +1
The pursuit of robustness has recently been a popular topic in reinforcement learning (RL) research, yet the existing methods generally suffer from efficiency issues that obstruct…