activity
20162026
most citedDegenerate Feedback Loops in Recommender Systems

153 citations · 362 across the 31 of their papers we have counts for

collaborators
Showing 2021Show all

6 papers · 1 filter

stat.ML2021

Variational Bayesian Optimistic Sampling

Brendan O'Donoghue, Tor Lattimore

We consider online sequential decision problems where an agent must balance exploration and exploitation. We derive a set of Bayesian `optimistic' policies which, in the stochastic…

cs.LG20211 cited

Minimax Regret for Bandit Convex Optimisation of Ridge Functions

Tor Lattimore

We analyse adversarial bandit convex optimisation with an adversary that is restricted to playing functions of the form for convex $g_t : \mathb…

stat.ML20213 cited

Bandit Phase Retrieval

Tor Lattimore, Botao Hao

We study a bandit version of phase retrieval where the learner chooses actions in the -dimensional unit ball and the expected reward is $\langle A_t, θ_\star\ran…

stat.ML20213 cited

Information Directed Sampling for Sparse Linear Bandits

Botao Hao, Tor Lattimore, Wei Deng

Stochastic sparse linear bandits offer a practical model for high-dimensional online decision-making problems and have a rich information-regret structure. In this work we explore…

cs.LG20212 cited

On the Optimality of Batch Policy Optimization Algorithms

Chenjun Xiao, Yifan Wu, Tor Lattimore +5

Batch policy optimization considers leveraging existing data for policy construction before interacting with an environment. Although interest in this problem has grown significant…

cs.LG20218 cited

Geometric Entropic Exploration

Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Alaa Saade +7

Exploration is essential for solving complex Reinforcement Learning (RL) tasks. Maximum State-Visitation Entropy (MSVE) formulates the exploration problem as a well-defined policy…