3 papers
cs.LG2026
VIPO: Value Function Inconsistency Penalized Offline Reinforcement Learning
Xuyang Chen, Keyu Yan, Guojian Wang +1
Offline reinforcement learning (RL) learns effective policies from pre-collected datasets, offering a practical solution for applications where online interactions are risky or cos…
cs.LG2026
One-Step Sampler for Boltzmann Distributions via Drifting
Wenhan Cao, Keyu Yan, Lin Zhao
We present a drifting-based framework for amortized sampling of Boltzmann distributions defined by energy functions. The method trains a one-step neural generator by projecting sam…
cs.LG2025
Taming OOD Actions for Offline Reinforcement Learning: An Advantage-Based Approach
Xuyang Chen, Keyu Yan, Wenhan Cao +1
Offline reinforcement learning (RL) learns policies from fixed datasets without online interactions, but suffers from distribution shift, causing inaccurate evaluation and overesti…