5 papers
Robust Offline Reinforcement Learning with Linearly Structured f-Divergence Regularization
Cheng Tang, Zhishuai Liu, Pan Xu
The Robust Regularized Markov Decision Process (RRMDP) is proposed to learn policies robust to dynamics shifts by adding regularization to the transition dynamics in the value func…
Randomized Exploration in Cooperative Multi-Agent Reinforcement Learning
Hao-Lun Hsu, Weixin Wang, Miroslav Pajic +1
We present the first study on provably efficient randomized exploration in cooperative multi-agent reinforcement learning (MARL). We propose a unified algorithm framework for rando…
Minimax Optimal and Computationally Efficient Algorithms for Distributionally Robust Offline Reinforcement Learning
Zhishuai Liu, Pan Xu
Distributionally robust offline reinforcement learning (RL), which seeks robust policy training against environment perturbation by modeling dynamics uncertainty, calls for functio…
Optimal Batched Best Arm Identification
Tianyuan Jin, Yu Yang, Jing Tang +2
We study the batched best arm identification (BBAI) problem, where the learner's goal is to identify the best arm while switching the policy as less as possible. In particular, we…
Off-Dynamics Reinforcement Learning via Domain Adaptation and Reward Augmented Imitation
Yihong Guo, Yixuan Wang, Yuanyuan Shi +2
Training a policy in a source domain for deployment in the target domain under a dynamics shift can be challenging, often resulting in performance degradation. Previous work tackle…