3 papers
cs.LG2026
AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL
Woosung Kim, Youngjun Suh, Jinho Lee +2
Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk aversion or fairness. However, such non-li…
cs.LG2026
Learning Policy from a Single Trajectory in Average-Reward Markov Decision Process
Jongmin Lee, Ernest K. Ryu, Vaneet Aggarwal
While there is an extensive body of work characterizing the sample complexity of discounted cumulative-reward MDPs, finite sample analyses for average-reward MDPs have been limited…
cs.LG2025
Semi-gradient DICE for Offline Constrained Reinforcement Learning
Woosung Kim, JunHo Seo, Jongmin Lee +1
Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliabl…