153 citations · 381 across the 39 of their papers we have counts for
9 papers · 2 filters
Adaptive Exploration in Linear Contextual Bandit
Botao Hao, Tor Lattimore, Csaba Szepesvari
Contextual bandits serve as a fundamental model for many sequential decision making tasks. The most popular theoretically justified approaches are based on the optimism principle.…
Gated Linear Networks
Joel Veness, Tor Lattimore, David Budden +8
This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distri…
Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel +11
This paper introduces the Behaviour Suite for Reinforcement Learning, or bsuite for short. bsuite is a collection of carefully-designed experiments that investigate core capabiliti…
Exploration by Optimisation in Partial Monitoring
Tor Lattimore, Csaba Szepesvari
We provide a simple and efficient algorithm for adversarial -action -outcome non-degenerate locally observable partial monitoring game for which the -round minimax regret…
Connections Between Mirror Descent, Thompson Sampling and the Information Ratio
Julian Zimmert, Tor Lattimore
The information-theoretic analysis by Russo and Van Roy (2014) in combination with minimax duality has proved a powerful tool for the analysis of online learning algorithms in full…
On First-Order Bounds, Variance and Gap-Dependent Bounds for Adversarial Bandits
Roman Pogodin, Tor Lattimore
We make three contributions to the theory of k-armed adversarial bandits. First, we prove a first-order bound for a modified variant of the INF strategy by Audibert and Bubeck [200…