Showing stat.MLShow all
3 papers · 1 filter
stat.ML2026
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
Kyungbok Lee, Angelica Cristello Sarteau, Michael R. Kosorok
We introduce a novel cyclic Markov decision process (MDP) framework for multi-step decision problems with heterogeneous stage-specific dynamics, transitions, and discount factors a…
stat.ML2024
Doubly-Robust Off-Policy Evaluation with Estimated Logging Policy
Kyungbok Lee, Myunghee Cho Paik
We introduce a novel doubly-robust (DR) off-policy evaluation (OPE) estimator for Markov decision processes, DRUnknown, designed for situations where both the logging policy and th…
stat.ML2023
Wasserstein Geodesic Generator for Conditional Distributions
Young-geun Kim, Kyungbok Lee, Youngwon Choi +2
Generating samples given a specific label requires estimating conditional distributions. We derive a tractable upper bound of the Wasserstein distance between conditional distribut…