2 papers
cs.LG2026
Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning
Sungyoung Lee, Dohyeong Kim, Eshan Balachandar +2
We propose Flow-Anchored Noise-conditioned Q-Learning (FAN), a highly efficient and high-performing offline reinforcement learning (RL) algorithm. Recent work has shown that expres…
cs.LG2026
Optimize Wider, Not Deeper: Consensus Aggregation for Policy Optimization
Zelal Su, Mustafaoglu, Sungyoung Lee +3
Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, crea…