activity
20192023
most citedSparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

10 citations · 33 across the 8 of their papers we have counts for

collaborators
Showing stat.MLShow all

7 papers · 1 filter

stat.ML20225 cited

Interacting Contour Stochastic Gradient Langevin Dynamics

Wei Deng, Siqi Liang, Botao Hao +2

We propose an interacting contour stochastic gradient Langevin dynamics (ICSGLD) sampler, an embarrassingly parallel multiple-chain contour stochastic gradient Langevin dynamics (C…

stat.ML20213 cited

Bandit Phase Retrieval

Tor Lattimore, Botao Hao

We study a bandit version of phase retrieval where the learner chooses actions in the -dimensional unit ball and the expected reward is $\langle A_t, θ_\star\ran…

stat.ML20213 cited

Information Directed Sampling for Sparse Linear Bandits

Botao Hao, Tor Lattimore, Wei Deng

Stochastic sparse linear bandits offer a practical model for high-dimensional online decision-making problems and have a rich information-regret structure. In this work we explore…

stat.ML2020

High-Dimensional Sparse Linear Bandits

Botao Hao, Tor Lattimore, Mengdi Wang

Stochastic linear bandits with high-dimensional sparse features are a practical model for a variety of domains, including personalized medicine and online advertising. We derive a…

stat.ML20205 cited

Residual Bootstrap Exploration for Bandit Algorithms

Chi-Hua Wang, Yang Yu, Botao Hao +1

In this paper, we propose a novel perturbation-based exploration method in bandit algorithms with bounded or unbounded rewards, called residual bootstrap exploration (\texttt{ReBoo…

stat.ML2019

Bootstrapping Upper Confidence Bound

Botao Hao, Yasin Abbasi-Yadkori, Zheng Wen +1

Upper Confidence Bound (UCB) method is arguably the most celebrated one used in online decision making with partial information feedback. Existing techniques for constructing confi…