2 papers
cs.LG2025
Semi-gradient DICE for Offline Constrained Reinforcement Learning
Woosung Kim, JunHo Seo, Jongmin Lee +1
Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliabl…
cs.LG2025
NBDI: A Simple and Effective Termination Condition for Skill Extraction from Task-Agnostic Demonstrations
Myunsoo Kim, Hayeong Lee, Seong-Woong Shim +2
Intelligent agents are able to make decisions based on different levels of granularity and duration. Recent advances in skill learning enabled the agent to solve complex, long-hori…