9 citations · 23 across the 4 of their papers we have counts for
4 papers
Towards Scalable and Robust Structured Bandits: A Meta-Learning Framework
Runzhe Wan, Lin Ge, Rui Song
Online learning in large-scale structured bandits is known to be challenging due to the curse of dimensionality. In this paper, we propose a unified meta-learning framework for a g…
Metadata-based Multi-Task Bandits with Bayesian Hierarchical Models
Runzhe Wan, Lin Ge, Rui Song
How to explore efficiently is a central problem in multi-armed bandits. In this paper, we introduce the metadata-based multi-task bandit problem, where the agent needs to solve a l…
Deeply-Debiased Off-Policy Interval Estimation
Chengchun Shi, Runzhe Wan, Victor Chernozhukov +1
Off-policy evaluation learns a target policy's value with a historical dataset generated by a different behavior policy. In addition to a point estimate, many applications would be…
Does the Markov Decision Process Fit the Data: Testing for the Markov Property in Sequential Decision Making
Chengchun Shi, Runzhe Wan, Rui Song +2
The Markov assumption (MA) is fundamental to the empirical validity of reinforcement learning. In this paper, we propose a novel Forward-Backward Learning procedure to test MA in s…