Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Semi-gradient DICE for Offline Constrained Reinforcement Learning
Woosung Kim, JunHo Seo, Jongmin Lee +1
Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliabl…
cs.LG2024
Offline Imitation Learning by Controlling the Effective Planning Horizon
Hee-Jun Ahn, Seong-Woong Shim, Byung-Jun Lee
In offline imitation learning (IL), we generally assume only a handful of expert trajectories and a supplementary offline dataset from suboptimal behaviors to learn the expert poli…