Publications (68)
StarCraft II: A New Challenge for Reinforcement Learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov +22
Value Iteration with Options and State Aggregation
Kamil Ciosek, David Silver
Reinforcement Learning via AIXI Approximation
Joel Veness, Kee Siong Ng, Marcus Hutter +1
Self-Consistent Models and Values
Gregory Farquhar, Kate Baumli, Zita Marinho +4
The Option Keyboard: Combining Skills in Reinforcement Learning
André Barreto, Diana Borsa, Shaobo Hou +8
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel +5
Meta-Gradient Reinforcement Learning with an Objective Discovered Online
Zhongwen Xu, Hado van Hasselt, Matteo Hessel +3
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
David Silver, Thomas Hubert, Julian Schrittwieser +10
Bootstrapped Meta-Learning
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy +3
Imagination-Augmented Agents for Deep Reinforcement Learning
Théophane Weber, Sébastien Racanière, David P. Reichert +12
Massively Parallel Methods for Deep Reinforcement Learning
Arun Nair, Praveen Srinivasan, Sam Blackwell +11
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning
Will Dabney, André Barreto, Mark Rowland +4
Meta-Gradient Reinforcement Learning
Zhongwen Xu, Hado van Hasselt, David Silver
Rainbow: Combining Improvements in Deep Reinforcement Learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt +7
A Monte Carlo AIXI Approximation
Joel Veness, Kee Siong Ng, Marcus Hutter +2
DataRater: Meta-Learned Dataset Curation
Dan A. Calian, Gregory Farquhar, Iurii Kemaev +9
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
The Predictron: End-To-End Learning and Planning
David Silver, Hado van Hasselt, Matteo Hessel +8
Prioritized Experience Replay
Tom Schaul, John Quan, Ioannis Antonoglou +1
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap +1
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor +10
The Value Equivalence Principle for Model-Based Reinforcement Learning
Christopher Grimm, André Barreto, Satinder Singh +1
Learning to Search with MCTSnets
Arthur Guez, Théophane Weber, Ioannis Antonoglou +5
Implicit Quantile Networks for Distributional Reinforcement Learning
Will Dabney, Georg Ostrovski, David Silver +1
Learning values across many orders of magnitude
Hado van Hasselt, Arthur Guez, Matteo Hessel +2
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert +9
Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement
André Barreto, Diana Borsa, John Quan +6
Discovery of Options via Meta-Learned Subgoals
Vivek Veeriah, Tom Zahavy, Matteo Hessel +6
Online and Offline Reinforcement Learning by Planning with a Learned Model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane +3
Proper Value Equivalence
Christopher Grimm, André Barreto, Gregory Farquhar +2
Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
Julien Perolat, Bart de Vylder, Daniel Hennes +31
Behaviour Suite for Reinforcement Learning
Ian Osband, Yotam Doron, Matteo Hessel +11
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
Human-level performance in first-person multiplayer games with population-based deep reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning +15
Muesli: Combining Improvements in Policy Optimization
Matteo Hessel, Ivo Danihelka, Fabio Viola +6
A Self-Tuning Actor-Critic Algorithm
Tom Zahavy, Zhongwen Xu, Vivek Veeriah +5
Reinforcement Learning with Unsupervised Auxiliary Tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki +4
Unicorn: Continual Learning with a Universal, Off-policy Agent
Daniel J. Mankowitz, Augustin ŽÃdek, André Barreto +7
FeUdal Networks for Hierarchical Reinforcement Learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul +4
Deep Reinforcement Learning with Double Q-learning
Hado van Hasselt, Arthur Guez, David Silver
Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
Johannes Heinrich, David Silver
Decoupled Neural Interfaces using Synthetic Gradients
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero +4
Better Optimism By Bayes: Adaptive Planning with Rich Models
Arthur Guez, David Silver, Peter Dayan
Learning to Win by Reading Manuals in a Monte-Carlo Framework
S. R. K. Branavan, David Silver, Regina Barzilay
Unit Tests for Stochastic Optimization
Tom Schaul, Ioannis Antonoglou, David Silver
Discovering Reinforcement Learning Algorithms
Junhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki +4
Unsupervised Predictive Memory in a Goal-Directed Agent
Greg Wayne, Chia-Chun Hung, David Amos +21
A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys +5
Learning Continuous Control Policies by Stochastic Value Gradients
Nicolas Heess, Greg Wayne, David Silver +3
Emergence of Locomotion Behaviours in Rich Environments
Nicolas Heess, Dhruva TB, Srinivasan Sriram +9
Successor Features for Transfer in Reinforcement Learning
André Barreto, Will Dabney, Rémi Munos +4
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
Learning and Transfer of Modulated Locomotor Controllers
Nicolas Heess, Greg Wayne, Yuval Tassa +3
Distributed Prioritized Experience Replay
Dan Horgan, John Quan, David Budden +4
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza +5
Efficient Bayes-Adaptive Reinforcement Learning using Sample-Based Search
Arthur Guez, David Silver, Peter Dayan
Discovery of Useful Questions as Auxiliary Tasks
Vivek Veeriah, Matteo Hessel, Zhongwen Xu +6
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver +4
Learning and Planning in Complex Action Spaces
Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou +3
Bayesian Optimization in AlphaGo
Yutian Chen, Aja Huang, Ziyu Wang +4
Universal Successor Features Approximators
Diana Borsa, André Barreto, John Quan +5
Expected Eligibility Traces
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel +3
Compositional Planning Using Optimal Option Models
David Silver, Kamil Ciosek
Credit Assignment Techniques in Stochastic Computation Graphs
Théophane Weber, Nicolas Heess, Lars Buesing +1
On Inductive Biases in Deep Reinforcement Learning
Matteo Hessel, Hado van Hasselt, Joseph Modayil +1
Move Evaluation in Go Using Deep Convolutional Neural Networks
Chris J. Maddison, Aja Huang, Ilya Sutskever +1
What Can Learned Intrinsic Rewards Capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel +5
Value-driven Hindsight Modelling
Arthur Guez, Fabio Viola, Théophane Weber +5