Publications (16)
Beyond Fine-Tuning: Transferring Behavior in Reinforcement Learning
VÃctor Campos, Pablo Sprechmann, Steven Hansen +5
Designing agents that acquire knowledge autonomously and use it to solve new tasks efficiently is an important challenge in reinforcement learning. Knowledge acquired during an uns…
Relative Variational Intrinsic Control
Kate Baumli, David Warde-Farley, Steven Hansen +1
In the absence of external rewards, agents can still learn useful behaviors by identifying and mastering a set of diverse skills within their environment. Existing skill learning m…
Fast deep reinforcement learning using online adjustments from the past
Steven Hansen, Pablo Sprechmann, Alexander Pritzel +2
We propose Ephemeral Value Adjusments (EVA): a means of allowing deep reinforcement learning agents to rapidly adapt to experience in their replay buffer. EVA shifts the value pred…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…
Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
Gemini Robotics Team, Abbas Abdolmaleki, Saminda Abeyruwan +169
General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the G…
Agency Is Frame-Dependent
David Abel, André Barreto, Michael Bowling +13
Agency is a system's capacity to steer outcomes toward a goal, and is a central topic of study across biology, philosophy, cognitive science, and artificial intelligence. Determini…
Generalization of Reinforcement Learners with Working and Episodic Memory
Meire Fortunato, Melissa Tan, Ryan Faulkner +6
Memory is an important aspect of intelligence and plays a role in many deep reinforcement learning models. However, little progress has been made in understanding when specific mem…
Unsupervised Control Through Non-Parametric Discriminative Rewards
David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni +3
Learning to control an environment without hand-crafted rewards or expert data remains challenging and is at the frontier of reinforcement learning research. We present an unsuperv…
Fast Task Inference with Variational Intrinsic Successor Features
Steven Hansen, Will Dabney, Andre Barreto +3
It has been established that diverse behaviors spanning the controllable subspace of an Markov decision process can be trained by rewarding a policy for being distinguishable from…
Plasticity as the Mirror of Empowerment
David Abel, Michael Bowling, André Barreto +13
Agents are minimally entities that are influenced by their past observations and act to influence future observations. This latter capacity is captured by empowerment, which has se…
Learning more skills through optimistic exploration
DJ Strouse, Kate Baumli, David Warde-Farley +2
Unsupervised skill learning objectives (Gregor et al., 2016, Eysenbach et al., 2018) allow agents to learn rich repertoires of behavior in the absence of extrinsic rewards. They wo…
Wasserstein Distance Maximizing Intrinsic Control
Ishan Durugkar, Steven Hansen, Stephen Spencer +1
This paper deals with the problem of learning a skill-conditioned policy that acts meaningfully in the absence of a reward signal. Mutual information based objectives have shown so…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
Uniqueness and Complexity of Inverse MDP Models
Marcus Hutter, Steven Hansen
What is the action sequence aa'a" that was likely responsible for reaching state s"' (from state s) in 3 steps? Addressing such questions is important in causal reasoning and in re…
In-context Reinforcement Learning with Algorithm Distillation
Michael Laskin, Luyu Wang, Junhyuk Oh +11
We propose Algorithm Distillation (AD), a method for distilling reinforcement learning (RL) algorithms into neural networks by modeling their training histories with a causal seque…