11 papers
Lipschitz Regularity in Wasserstein Robust Stochastic Optimal Control
Shengbo Wang, Jose Blanchet
Robust Markov decision processes provide a principled framework for protecting sequential decision-making against transition-law misspecification and have attracted substantial rec…
Fast Convergence of Policy Regret in Learning Stochastic Optimal Control
Shengbo Wang, Jose Blanchet, Peter Glynn
Policy learning in modern operations environments faces a fundamental tension between limited operational data and the large, often continuous, state and action spaces over which g…
Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
Shengbo Wang, Zexi Zhang
Designing model-free algorithms for distributionally robust reinforcement learning (DRRL) poses fundamental challenges. The robust Bellman operator is nonlinear in the transition k…
Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning
Zhenghao Li, Shengbo Wang, Nian Si
Distributionally robust reinforcement learning (DR-RL) has recently gained significant attention as a principled approach that addresses discrepancies between training and testing…
Non-Rectangular Average-Reward Robust MDPs: Optimal Policies and Their Transient Values
Shengbo Wang, Nian Si
We study non-rectangular robust Markov decision processes under the average-reward criterion, where the ambiguity set couples transition probabilities across states and the adversa…
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
Zijun Chen, Shengbo Wang, Nian Si
Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally ro…