Publications (13)
Vision-Language Models as a Source of Rewards
Kate Baumli, Satinder Baveja, Feryal Behbahani +24
Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for bui…
Gemini: A Family of Highly Capable Multimodal Models
Gemini Team, Rohan Anil, Sebastian Borgeaud +1340
This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…
Predicting Ordinary Differential Equations with Transformers
Sören Becker, Michal Klein, Alexander Neitz +2
We develop a transformer-based sequence-to-sequence model that recovers scalar ordinary differential equations (ODEs) in symbolic form from irregularly sampled and noisy observatio…
Discovering ordinary differential equations that govern time-series
Sören Becker, Michal Klein, Alexander Neitz +2
Natural laws are often described through differential equations yet finding a differential equation that describes the governing law underlying observed data is a challenging and s…
Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning
Giambattista Parascandolo, Lars Buesing, Josh Merel +6
Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumpt…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…
Learning explanations that are hard to vary
Giambattista Parascandolo, Alexander Neitz, Antonio Orvieto +2
In this paper, we investigate the principle that `good explanations are hard to vary' in the context of deep learning. We show that averaging gradients across examples -- akin to a…
Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models
Alexander Neitz, Giambattista Parascandolo, Stefan Bauer +1
We introduce a method which enables a recurrent dynamics model to be temporally abstract. Our approach, which we call Adaptive Skip Intervals (ASI), is based on the observation tha…
Direct Advantage Estimation
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz +1
The predominant approach in reinforcement learning is to assign credit to actions based on the expected return. However, we show that the return may depend on the policy in a way w…
Neural Symbolic Regression that Scales
Luca Biggio, Tommaso Bendinelli, Alexander Neitz +2
Symbolic equations are at the core of scientific discovery. The task of discovering the underlying equation from a set of input-output pairs is called symbolic regression. Traditio…
OpenAI o1 System Card
OpenAI, :, Aaron Jaech +261
The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the…
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI, :, Sandhini Agarwal +124
We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…
CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning
Ossama Ahmed, Frederik Träuble, Anirudh Goyal +5
Despite recent successes of reinforcement learning (RL), it remains a challenge for agents to transfer learned skills to related environments. To facilitate research addressing thi…