13 citations · 34 across the 16 of their papers we have counts for
6 papers · 1 filter
Variational Bayesian Optimistic Sampling
Brendan O'Donoghue, Tor Lattimore
We consider online sequential decision problems where an agent must balance exploration and exploitation. We derive a set of Bayesian `optimistic' policies which, in the stochastic…
The Neural Testbed: Evaluating Joint Predictions
Ian Osband, Zheng Wen, Seyed Mohammad Asghari +7
Predictive distributions quantify uncertainties ignored by point estimates. This paper introduces The Neural Testbed: an open-source benchmark for controlled and principled evaluat…
Practical Large-Scale Linear Programming using Primal-Dual Hybrid Gradient
David Applegate, Mateo Díaz, Oliver Hinder +4
We present PDLP, a practical first-order method for linear programming (LP) that can solve to the high levels of accuracy that are expected in traditional LP applications. In addit…
Discovering Diverse Nearly Optimal Policies with Successor Features
Tom Zahavy, Brendan O'Donoghue, Andre Barreto +3
Finding different solutions to the same problem is a key aspect of intelligence associated with creativity and adaptation to novel situations. In reinforcement learning, a set of d…
Reward is enough for convex MDPs
Tom Zahavy, Brendan O'Donoghue, Guillaume Desjardins +1
Maximising a cumulative reward function that is Markov and stationary, i.e., defined over state-action pairs and independent of time, is sufficient to capture many kinds of goals i…
Discovering a set of policies for the worst case reward
Tom Zahavy, Andre Barreto, Daniel J Mankowitz +4
We study the problem of how to construct a set of policies that can be composed together to solve a collection of reinforcement learning tasks. Each task is a different reward func…