3 papers
cs.LG2023
Minimax Weight Learning for Absorbing MDPs
Fengyin Li, Yuqiang Li, Xianyi Wu
Reinforcement learning policy evaluation problems are often modeled as finite or discounted/averaged infinite-horizon MDPs. In this paper, we study undiscounted off-policy policy e…
math.PR2018
A General Framework of Multi-Armed Bandit Processes by Arm Switch Restrictions
Wenqing Bao, Xiaoqiang Cai, Xianyi Wu
This paper proposes a general framework of multi-armed bandit (MAB) processes by introducing a type of restrictions on the switches among arms evolving in continuous time. The Gitt…
math.ST2016
On Hodges' Superefficiency and Merits of Oracle Property in Model Selection
Xianyi Wu, Xian Zhou
The oracle property of model selection procedures has attracted a large volume of favorable publications in the literature, but also faced criticisms of being ineffective and misle…