papers

Publications (39)

cs.RO2025

Gemini Robotics: Bringing AI into the Physical World

Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie +115

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as…

cs.CV2024

A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…

cs.CV2014

OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks

Pierre Sermanet, David Eigen, Xiang Zhang +3

We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be effi…

cs.RO2023

Visuomotor Control in Multi-Object Scenes Using Object-Aware Representations

Negin Heravi, Ayzaan Wahid, Corey Lynch +6

Perceptual understanding of the scene and the relationship between its different components is important for successful completion of robotic tasks. Representation learning has bee…

cs.CV2020

Counting Out Time: Class Agnostic Video Repetition Counting in the Wild

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temp…

cs.RO2020

Motion2Vec: Semi-Supervised Representation Learning from Surgical Videos

Ajay Kumar Tanwani, Pierre Sermanet, Andy Yan +3

Learning meaningful visual representations in an embedding space can facilitate generalization in downstream tasks such as action segmentation and imitation. In this paper, we lear…

cs.CV2019

Temporal Cycle-Consistency Learning

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

We introduce a self-supervised representation learning method based on the task of temporal alignment between videos. The method trains a network using temporal cycle consistency (…

cs.RO2024

RT-H: Action Hierarchies Using Language

Suneel Belkhale, Tianli Ding, Ted Xiao +6

Language provides a way to break down complex concepts into digestible pieces. Recent works in robot imitation learning use language-conditioned policies that predict actions given…

cs.RO2023

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, Noah Brown, Justice Carbajal +51

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic…

cs.LG2019

Wasserstein Dependency Measure for Representation Learning

Sherjil Ozair, Corey Lynch, Yoshua Bengio +3

Mutual information maximization has emerged as a powerful learning objective for unsupervised representation learning obtaining state-of-the-art performance in applications such as…

cs.CV2015

Attention for Fine-Grained Categorization

Pierre Sermanet, Andrea Frome, Esteban Real

This paper presents experiments extending the work of Ba et al. (2014) on recurrent neural models for attention into less constrained visual environments, specifically fine-grained…

cs.CL2025

SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior

Pierre Sermanet, Anirudha Majumdar, Vikas Sindhwani

Given the recent rate of progress in artificial intelligence (AI) and robotics, a tantalizing question is emerging: would robots controlled by emerging AI systems be strongly align…

cs.RO2025

Generating Robot Constitutions & Benchmarks for Semantic Safety

Pierre Sermanet, Anirudha Majumdar, Alex Irpan +2

Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and…

cs.RO2024

Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Vidhi Jain, Maria Attarian, Nikhil J Joshi +10

Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the ro…

cs.RO2020

Broadly-Exploring, Local-Policy Trees for Long-Horizon Task Planning

Brian Ichter, Pierre Sermanet, Corey Lynch

Long-horizon planning in realistic environments requires the ability to reason over sequential tasks in high-dimensional state spaces with complex dynamics. Classical motion planni…

cs.RO2025

Robotic Table Tennis: A Case Study into a High Speed Learning System

David B. D'Ambrosio, Jonathan Abelian, Saminda Abeyruwan +32

We present a deep-dive into a real-world robotic learning system that, in previous work, was shown to be capable of hundreds of table tennis rallies with a human and has the abilit…

cs.CV2019

Online Object Representations with Contrastive Learning

Sören Pirk, Mohi Khansari, Yunfei Bai +2

We propose a self-supervised approach for learning representations of objects from monocular videos and demonstrate it is particularly useful in situated settings such as robotics.…

cs.RO2022

GoalsEye: Learning High Speed Precision Table Tennis on a Physical Robot

Tianli Ding, Laura Graesser, Saminda Abeyruwan +5

Learning goal conditioned control in the real world is a challenging open problem in robotics. Reinforcement learning systems have the potential to learn autonomously via trial-and…

cs.CV2012

Convolutional Neural Networks Applied to House Numbers Digit Classification

Pierre Sermanet, Soumith Chintala, Yann LeCun

We classify digits of real-world house numbers using convolutional neural networks (ConvNets). ConvNets are hierarchical feature learning neural networks whose structure is biologi…

cs.RO2019

Learning Latent Plans from Play

Corey Lynch, Mohi Khansari, Ted Xiao +4

Acquiring a diverse repertoire of general-purpose skills remains an open challenge for robotics. In this work, we propose self-supervising control on top of human teleoperated play…

cs.CV2013

Pedestrian Detection with Unsupervised Multi-Stage Feature Learning

Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala +1

Pedestrian detection is a problem of considerable practical interest. Adding to the list of successful applications of deep learning methods to vision, we report state-of-the-art a…

cs.LG2023

PaLM-E: An Embodied Multimodal Language Model

Danny Driess, Fei Xia, Mehdi S. M. Sajjadi +19

Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding.…

cs.CV2018

Time-Contrastive Networks: Self-Supervised Learning from Video

Pierre Sermanet, Corey Lynch, Yevgen Chebotar +4

We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this repres…

cs.CL2025

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431

In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…

cs.CV2014

Going Deeper with Convolutions

Christian Szegedy, Wei Liu, Yangqing Jia +6

We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in th…

cs.CV2021

With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…

cs.RO2024

AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

Michael Ahn, Debidatta Dwibedi, Chelsea Finn +24

Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However,…

cs.CV2023

Video Language Planning

Yilun Du, Mengjiao Yang, Pete Florence +10

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…

cs.RO2025

Predictive Red Teaming: Breaking Policies Without Breaking Robots

Anirudha Majumdar, Mohit Sharma, Dmitry Kalashnikov +3

Visuomotor policies trained via imitation learning are capable of performing challenging manipulation tasks, but are often extremely brittle to lighting, visual distractors, and ob…

cs.RO2025

Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…

cs.RO2023

RoboVQA: Multimodal Long-Horizon Reasoning for Robotics

Pierre Sermanet, Tianli Ding, Jeffrey Zhao +18

We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher t…

cs.RO2022

Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Michael Ahn, Anthony Brohan, Noah Brown +42

Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extend…

cs.RO2022

Inner Monologue: Embodied Reasoning through Planning with Language Models

Wenlong Huang, Fei Xia, Ted Xiao +14

Recent works have shown how the reasoning capabilities of Large Language Models (LLMs) can be applied to domains beyond natural language processing, such as planning and interactio…

cs.RO2023

Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models

Ted Xiao, Harris Chan, Pierre Sermanet +5

In recent years, much progress has been made in learning robotic manipulation policies that follow natural language instructions. Such methods typically learn from corpora of robot…

cs.CV2019

Learning Actionable Representations from Visual Observations

Debidatta Dwibedi, Jonathan Tompson, Corey Lynch +1

In this work we explore a new approach for robots to teach themselves about the world simply by observing it. In particular we investigate the effectiveness of learning task-agnost…

cs.CV2017

Unsupervised Perceptual Rewards for Imitation Learning

Pierre Sermanet, Kelvin Xu, Sergey Levine

Reward function design and exploration time are arguably the biggest obstacles to the deployment of reinforcement learning (RL) agents in the real world. In many real-world tasks,…

cs.RO2021

Language Conditioned Imitation Learning over Unstructured Data

Corey Lynch, Pierre Sermanet

Natural language is perhaps the most flexible and intuitive way for humans to communicate tasks to a robot. Prior work in imitation learning typically requires each task be specifi…

cs.AI2025

Can AI Perceive Physical Danger and Intervene?

Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer +5

When AI interacts with the physical world -- as a robot or an assistive agent -- new safety challenges emerge beyond those of purely ``digital AI". In such interactions, the potent…

cs.RO2020

Learning to Play by Imitating Humans

Rostam Dinyari, Pierre Sermanet, Corey Lynch

Acquiring multiple skills has commonly involved collecting a large number of expert demonstrations per task or engineering custom reward functions. Recently it has been shown that…