papers

Publications (13)

cs.LG2024

Vision-Language Models as a Source of Rewards

Kate Baumli, Satinder Baveja, Feryal Behbahani +24

Building generalist agents that can accomplish many goals in rich open-ended environments is one of the research frontiers for reinforcement learning. A key limiting factor for bui…

cs.CL2025

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +1340

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consist…

cs.LG2023

Predicting Ordinary Differential Equations with Transformers

Sören Becker, Michal Klein, Alexander Neitz +2

We develop a transformer-based sequence-to-sequence model that recovers scalar ordinary differential equations (ODEs) in symbolic form from irregularly sampled and noisy observatio…

cs.LG2022

Discovering ordinary differential equations that govern time-series

Sören Becker, Michal Klein, Alexander Neitz +2

Natural laws are often described through differential equations yet finding a differential equation that describes the governing law underlying observed data is a challenging and s…

cs.LG2020

Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning

Giambattista Parascandolo, Lars Buesing, Josh Merel +6

Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumpt…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.LG2020

Learning explanations that are hard to vary

Giambattista Parascandolo, Alexander Neitz, Antonio Orvieto +2

In this paper, we investigate the principle that `good explanations are hard to vary' in the context of deep learning. We show that averaging gradients across examples -- akin to a…

cs.LG2018

Adaptive Skip Intervals: Temporal Abstraction for Recurrent Dynamical Models

Alexander Neitz, Giambattista Parascandolo, Stefan Bauer +1

We introduce a method which enables a recurrent dynamics model to be temporally abstract. Our approach, which we call Adaptive Skip Intervals (ASI), is based on the observation tha…

cs.LG2023

Direct Advantage Estimation

Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz +1

The predominant approach in reinforcement learning is to assign credit to actions based on the expected return. However, we show that the return may depend on the policy in a way w…

cs.LG2021

Neural Symbolic Regression that Scales

Luca Biggio, Tommaso Bendinelli, Alexander Neitz +2

Symbolic equations are at the core of scientific discovery. The task of discovering the underlying equation from a set of input-output pairs is called symbolic regression. Traditio…

cs.AI2026

OpenAI o1 System Card

OpenAI, :, Aaron Jaech +261

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the…

cs.CL2025

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI, :, Sandhini Agarwal +124

We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…

cs.RO2020

CausalWorld: A Robotic Manipulation Benchmark for Causal Structure and Transfer Learning

Ossama Ahmed, Frederik Träuble, Anirudh Goyal +5

Despite recent successes of reinforcement learning (RL), it remains a challenge for agents to transfer learned skills to related environments. To facilitate research addressing thi…