Publications (13)
See the Glass Half Full: Reasoning about Liquid Containers, their Volume and Content
Roozbeh Mottaghi, Connor Schenck, Dieter Fox +1
Humans have rich understanding of liquid containers and their contents; for example, we can effortlessly pour water from a pitcher to a cup. Doing so requires estimating the volume…
CLAMP: Contrastive Learning for 3D Multi-View Action-Conditioned Robotic Manipulation Pretraining
I-Chun Arthur Liu, Krzysztof Choromanski, Sandy Huang +1
Leveraging pre-trained 2D image representations in behavior cloning policies has achieved great success and has become a standard approach for robotic manipulation. However, such r…
Reasoning About Liquids via Closed-Loop Simulation
Connor Schenck, Dieter Fox
Simulators are powerful tools for reasoning about a robot's interactions with its environment. However, when simulations diverge from reality, that reasoning becomes less useful. I…
Learning Robotic Manipulation of Granular Media
Connor Schenck, Jonathan Tompson, Dieter Fox +1
In this paper, we examine the problem of robotic manipulation of granular media. We evaluate multiple predictive models used to infer the dynamics of scooping and dumping actions.…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Towards Learning to Perceive and Reason About Liquids
Connor Schenck, Dieter Fox
Recent advances in AI and robotics have claimed many incredible results with deep learning, yet no work to date has applied deep learning to the problem of liquid perception and re…
Detection and Tracking of Liquids with Fully Convolutional Networks
Connor Schenck, Dieter Fox
Recent advances in AI and robotics have claimed many incredible results with deep learning, yet no work to date has applied deep learning to the problem of liquid perception and re…
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
Connor Schenck, Isaac Reid, Mithun George Jacob +19
We introduce STRING: Separable Translationally Invariant Position Encodings. STRING extends Rotary Position Encodings, a recently proposed and widely used algorithm in large langua…
Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu +19
An elusive goal in navigation research is to build an intelligent agent that can understand multimodal instructions including natural language and image, and perform useful navigat…
SPNets: Differentiable Fluid Dynamics for Deep Neural Networks
Connor Schenck, Dieter Fox
In this paper we introduce Smooth Particle Networks (SPNets), a framework for integrating fluid dynamics with deep networks. SPNets adds two new layers to the neural network toolbo…
Guided Policy Search with Delayed Sensor Measurements
Connor Schenck, Dieter Fox
Guided policy search is a method for reinforcement learning that trains a general policy for accomplishing a given task by guiding the learning of the policy with multiple guiding…
Visual Closed-Loop Control for Pouring Liquids
Connor Schenck, Dieter Fox
Pouring a specific amount of liquid is a challenging task. In this paper we develop methods for robots to use visual feedback to perform closed-loop control for pouring liquids. We…
Linear Transformer Topological Masking with Graph Random Features
Isaac Reid, Kumar Avinava Dubey, Deepali Jain +12
When training transformers on graph-structured data, incorporating information about the underlying topology is crucial for good performance. Topological masking, a type of relativ…