Showing math.OCShow all
3 papers · 1 filter
math.OC2026
Robust Markov Decision Processes on Continuous State Spaces
Mengmeng Li, Yifan Hu, Daniel Kuhn +1
We study infinite-horizon robust Markov decision processes (MDPs) on continuous state spaces with structured rectangular ambiguity set. The proposed ambiguity set falls within the…
math.OC2025
Towards Optimal Offline Reinforcement Learning
Mengmeng Li, Daniel Kuhn, Tobias Sutter
We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…
math.OC2024
A Large Deviations Perspective on Policy Gradient Algorithms
Wouter Jongeneel, Daniel Kuhn, Mengmeng Li
Motivated by policy gradient methods in the context of reinforcement learning, we identify a large deviation rate function for the iterates generated by stochastic gradient descent…