Publications (39)
Gemini Robotics: Bringing AI into the Physical World
Gemini Robotics Team, Saminda Abeyruwan, Joshua Ainslie +115
Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as…
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We discuss some consistent issues on how RepNet has been evaluated in various papers. As a way to mitigate these issues, we report RepNet performance results on different datasets,…
OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
Pierre Sermanet, David Eigen, Xiang Zhang +3
We present an integrated framework for using Convolutional Networks for classification, localization and detection. We show how a multiscale and sliding window approach can be effi…
Visuomotor Control in Multi-Object Scenes Using Object-Aware Representations
Negin Heravi, Ayzaan Wahid, Corey Lynch +6
Perceptual understanding of the scene and the relationship between its different components is important for successful completion of robotic tasks. Representation learning has bee…
Counting Out Time: Class Agnostic Video Repetition Counting in the Wild
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temp…
Motion2Vec: Semi-Supervised Representation Learning from Surgical Videos
Ajay Kumar Tanwani, Pierre Sermanet, Andy Yan +3
Learning meaningful visual representations in an embedding space can facilitate generalization in downstream tasks such as action segmentation and imitation. In this paper, we lear…
Temporal Cycle-Consistency Learning
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
We introduce a self-supervised representation learning method based on the task of temporal alignment between videos. The method trains a network using temporal cycle consistency (…
RT-H: Action Hierarchies Using Language
Suneel Belkhale, Tianli Ding, Ted Xiao +6
Language provides a way to break down complex concepts into digestible pieces. Recent works in robot imitation learning use language-conditioned policies that predict actions given…
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
Anthony Brohan, Noah Brown, Justice Carbajal +51
We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic…
Wasserstein Dependency Measure for Representation Learning
Sherjil Ozair, Corey Lynch, Yoshua Bengio +3
Mutual information maximization has emerged as a powerful learning objective for unsupervised representation learning obtaining state-of-the-art performance in applications such as…
Attention for Fine-Grained Categorization
Pierre Sermanet, Andrea Frome, Esteban Real
This paper presents experiments extending the work of Ba et al. (2014) on recurrent neural models for attention into less constrained visual environments, specifically fine-grained…
SciFi-Benchmark: Leveraging Science Fiction To Improve Robot Behavior
Pierre Sermanet, Anirudha Majumdar, Vikas Sindhwani
Given the recent rate of progress in artificial intelligence (AI) and robotics, a tantalizing question is emerging: would robots controlled by emerging AI systems be strongly align…
Generating Robot Constitutions & Benchmarks for Semantic Safety
Pierre Sermanet, Anirudha Majumdar, Alex Irpan +2
Until recently, robotics safety research was predominantly about collision avoidance and hazard reduction in the immediate vicinity of a robot. Since the advent of large vision and…
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
Vidhi Jain, Maria Attarian, Nikhil J Joshi +10
Large-scale multi-task robotic manipulation systems often rely on text to specify the task. In this work, we explore whether a robot can learn by observing humans. To do so, the ro…
Broadly-Exploring, Local-Policy Trees for Long-Horizon Task Planning
Brian Ichter, Pierre Sermanet, Corey Lynch
Long-horizon planning in realistic environments requires the ability to reason over sequential tasks in high-dimensional state spaces with complex dynamics. Classical motion planni…
Robotic Table Tennis: A Case Study into a High Speed Learning System
David B. D'Ambrosio, Jonathan Abelian, Saminda Abeyruwan +32
We present a deep-dive into a real-world robotic learning system that, in previous work, was shown to be capable of hundreds of table tennis rallies with a human and has the abilit…
Online Object Representations with Contrastive Learning
Sören Pirk, Mohi Khansari, Yunfei Bai +2
We propose a self-supervised approach for learning representations of objects from monocular videos and demonstrate it is particularly useful in situated settings such as robotics.…
GoalsEye: Learning High Speed Precision Table Tennis on a Physical Robot
Tianli Ding, Laura Graesser, Saminda Abeyruwan +5
Learning goal conditioned control in the real world is a challenging open problem in robotics. Reinforcement learning systems have the potential to learn autonomously via trial-and…
Convolutional Neural Networks Applied to House Numbers Digit Classification
Pierre Sermanet, Soumith Chintala, Yann LeCun
We classify digits of real-world house numbers using convolutional neural networks (ConvNets). ConvNets are hierarchical feature learning neural networks whose structure is biologi…
Learning Latent Plans from Play
Corey Lynch, Mohi Khansari, Ted Xiao +4
Acquiring a diverse repertoire of general-purpose skills remains an open challenge for robotics. In this work, we propose self-supervising control on top of human teleoperated play…
Pedestrian Detection with Unsupervised Multi-Stage Feature Learning
Pierre Sermanet, Koray Kavukcuoglu, Soumith Chintala +1
Pedestrian detection is a problem of considerable practical interest. Adding to the list of successful applications of deep learning methods to vision, we report state-of-the-art a…
PaLM-E: An Embodied Multimodal Language Model
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi +19
Large language models excel at a wide range of complex tasks. However, enabling general inference in the real world, e.g., for robotics problems, raises the challenge of grounding.…
Time-Contrastive Networks: Self-Supervised Learning from Video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar +4
We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this repres…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Going Deeper with Convolutions
Christian Szegedy, Wei Liu, Yangqing Jia +6
We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in th…
With a Little Help from My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations
Debidatta Dwibedi, Yusuf Aytar, Jonathan Tompson +2
Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat di…
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents
Michael Ahn, Debidatta Dwibedi, Chelsea Finn +24
Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However,…
Video Language Planning
Yilun Du, Mengjiao Yang, Pete Florence +10
We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pr…
Predictive Red Teaming: Breaking Policies Without Breaking Robots
Anirudha Majumdar, Mohit Sharma, Dmitry Kalashnikov +3
Visuomotor policies trained via imitation learning are capable of performing challenging manipulation tasks, but are often extremely brittle to lighting, visual distractors, and ob…
Open X-Embodiment: Robotic Learning Datasets and RT-X Models
Embodiment Collaboration, Abby O'Neill, Abdul Rehman +291
Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, thi…
RoboVQA: Multimodal Long-Horizon Reasoning for Robotics
Pierre Sermanet, Tianli Ding, Jeffrey Zhao +18
We present a scalable, bottom-up and intrinsically diverse data collection scheme that can be used for high-level reasoning with long and medium horizons and that has 2.2x higher t…
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Michael Ahn, Anthony Brohan, Noah Brown +42
Large language models can encode a wealth of semantic knowledge about the world. Such knowledge could be extremely useful to robots aiming to act upon high-level, temporally extend…
Inner Monologue: Embodied Reasoning through Planning with Language Models
Wenlong Huang, Fei Xia, Ted Xiao +14
Recent works have shown how the reasoning capabilities of Large Language Models (LLMs) can be applied to domains beyond natural language processing, such as planning and interactio…
Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models
Ted Xiao, Harris Chan, Pierre Sermanet +5
In recent years, much progress has been made in learning robotic manipulation policies that follow natural language instructions. Such methods typically learn from corpora of robot…
Learning Actionable Representations from Visual Observations
Debidatta Dwibedi, Jonathan Tompson, Corey Lynch +1
In this work we explore a new approach for robots to teach themselves about the world simply by observing it. In particular we investigate the effectiveness of learning task-agnost…
Unsupervised Perceptual Rewards for Imitation Learning
Pierre Sermanet, Kelvin Xu, Sergey Levine
Reward function design and exploration time are arguably the biggest obstacles to the deployment of reinforcement learning (RL) agents in the real world. In many real-world tasks,…
Language Conditioned Imitation Learning over Unstructured Data
Corey Lynch, Pierre Sermanet
Natural language is perhaps the most flexible and intuitive way for humans to communicate tasks to a robot. Prior work in imitation learning typically requires each task be specifi…
Can AI Perceive Physical Danger and Intervene?
Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer +5
When AI interacts with the physical world -- as a robot or an assistive agent -- new safety challenges emerge beyond those of purely ``digital AI". In such interactions, the potent…
Learning to Play by Imitating Humans
Rostam Dinyari, Pierre Sermanet, Corey Lynch
Acquiring multiple skills has commonly involved collecting a large number of expert demonstrations per task or engineering custom reward functions. Recently it has been shown that…