activity
20192022
most citedSparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

10 citations · 31 across the 7 of their papers we have counts for

collaborators

12 papers

stat.ML20225 cited

Interacting Contour Stochastic Gradient Langevin Dynamics

Wei Deng, Siqi Liang, Botao Hao +2

We propose an interacting contour stochastic gradient Langevin dynamics (ICSGLD) sampler, an embarrassingly parallel multiple-chain contour stochastic gradient Langevin dynamics (C…

stat.ML20213 cited

Bandit Phase Retrieval

Tor Lattimore, Botao Hao

We study a bandit version of phase retrieval where the learner chooses actions in the -dimensional unit ball and the expected reward is $\langle A_t, θ_\star\ran…

stat.ML20213 cited

Information Directed Sampling for Sparse Linear Bandits

Botao Hao, Tor Lattimore, Wei Deng

Stochastic sparse linear bandits offer a practical model for high-dimensional online decision-making problems and have a rich information-regret structure. In this work we explore…

cs.LG20212 cited

Optimization Issues in KL-Constrained Approximate Policy Iteration

Nevena Lazić, Botao Hao, Yasin Abbasi-Yadkori +2

Many reinforcement learning algorithms can be seen as versions of approximate policy iteration (API). While standard API often performs poorly, it has been shown that learning can…

cs.LG202010 cited

Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

Botao Hao, Yaqi Duan, Tor Lattimore +2

This paper provides a statistical analysis of high-dimensional batch Reinforcement Learning (RL) using sparse linear function approximation. When there is a large number of candida…

stat.ML2020

High-Dimensional Sparse Linear Bandits

Botao Hao, Tor Lattimore, Mengdi Wang

Stochastic linear bandits with high-dimensional sparse features are a practical model for a variety of domains, including personalized medicine and online advertising. We derive a…