22 citations · 28 across the 12 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
Diverse Policy Optimization for Structured Action Space
Wenhao Li, Baoxiang Wang, Shanchao Yang +1
Enhancing the diversity of policies is beneficial for robustness, exploration, and transfer in reinforcement learning (RL). In this paper, we aim to seek diverse policies in an und…
cs.LG2023
Improved Regret Bounds for Linear Adversarial MDPs via Linear Optimization
Fang Kong, Xiangcheng Zhang, Baoxiang Wang +1
Learning Markov decision processes (MDP) in an adversarial environment has been a challenging problem. The problem becomes even more challenging with function approximation, since…