153 citations · 362 across the 31 of their papers we have counts for
6 papers · 1 filter
Variational Bayesian Optimistic Sampling
Brendan O'Donoghue, Tor Lattimore
We consider online sequential decision problems where an agent must balance exploration and exploitation. We derive a set of Bayesian `optimistic' policies which, in the stochastic…
Minimax Regret for Bandit Convex Optimisation of Ridge Functions
Tor Lattimore
We analyse adversarial bandit convex optimisation with an adversary that is restricted to playing functions of the form for convex $g_t : \mathb…
Bandit Phase Retrieval
Tor Lattimore, Botao Hao
We study a bandit version of phase retrieval where the learner chooses actions in the -dimensional unit ball and the expected reward is $\langle A_t, θ_\star\ran…
Information Directed Sampling for Sparse Linear Bandits
Botao Hao, Tor Lattimore, Wei Deng
Stochastic sparse linear bandits offer a practical model for high-dimensional online decision-making problems and have a rich information-regret structure. In this work we explore…
On the Optimality of Batch Policy Optimization Algorithms
Chenjun Xiao, Yifan Wu, Tor Lattimore +5
Batch policy optimization considers leveraging existing data for policy construction before interacting with an environment. Although interest in this problem has grown significant…
Geometric Entropic Exploration
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Alaa Saade +7
Exploration is essential for solving complex Reinforcement Learning (RL) tasks. Maximum State-Visitation Entropy (MSVE) formulates the exploration problem as a well-defined policy…