Publications (17)
ViNT: A Foundation Model for Visual Navigation
Dhruv Shah, Ajay Sridhar, Nitish Dashora +4
General-purpose pre-trained models ("foundation models") have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that a…
: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
Physical Intelligence, Bo Ai, Ali Amin +85
We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language…
RACER: Epistemic Risk-Sensitive RL Enables Fast Driving with Fewer Crashes
Kyle Stachowicz, Sergey Levine
Reinforcement learning provides an appealing framework for robotic control due to its ability to learn expressive policies purely through real-world interaction. However, this requ…
Learning to Drive Anywhere with Model-Based Reannotation
Noriaki Hirose, Lydia Ignatova, Kyle Stachowicz +3
Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data.…
Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
Joshua Jones, Oier Mees, Carmelo Sferrazza +3
Interacting with the world is a multi-sensory experience: achieving effective general-purpose interaction requires making use of all available modalities -- including vision, touch…
Multimodal Maximum Entropy Dynamic Games
Oswin So, Kyle Stachowicz, Evangelos A. Theodorou
Environments with multi-agent interactions often result a rich set of modalities of behavior between agents due to the inherent suboptimality of decision making processes when agen…
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
Tian Gao, Celine Tan, Catherine Glossop +8
A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While l…
SELFI: Autonomous Self-Improvement with Reinforcement Learning for Social Navigation
Noriaki Hirose, Dhruv Shah, Kyle Stachowicz +2
Autonomous self-improving robots that interact and improve with experience are key to the real-world deployment of robotic systems. In this paper, we propose an online learning met…
MEM: Multi-Scale Embodied Memory for Vision Language Action Models
Marcel Torne, Karl Pertsch, Homer Walke +14
Conventionally, memory in end-to-end robotic learning involves inputting a sequence of past observations into the learned policy. However, in complex multi-stage real-world tasks,…
Traversability-Aware Legged Navigation by Learning from Real-World Visual Data
Hongbo Zhang, Zhongyu Li, Xuanqi Zeng +9
The enhanced mobility brought by legged locomotion empowers quadrupedal robots to navigate through complex and unstructured environments. However, optimizing agile locomotion while…
Safety Embedded Differential Dynamic Programming Using Discrete Barrier States
Hassan Almubarak, Kyle Stachowicz, Nader Sadegh +1
Certified safe control is a growing challenge in robotics, especially when performance and safety objectives must be concurrently achieved. In this work, we extend the barrier stat…
: a Vision-Language-Action Model with Open-World Generalization
Physical Intelligence, Kevin Black, Noah Brown +33
In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated im…
: a VLA That Learns From Experience
Physical Intelligence, Ali Amin, Raichelle Aniceto +53
We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience…
Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
Filippo Lazzati, Kyle Stachowicz, William Chen +3
Action chunking---predicting and executing multiple actions instead of a single action---has proven to be a critical component for learning effective robotic control policies. Howe…
Optimal-Horizon Model-Predictive Control with Differential Dynamic Programming
Kyle Stachowicz, Evangelos A. Theodorou
We present an algorithm, based on the Differential Dynamic Programming framework, to handle trajectory optimization problems in which the horizon is determined online rather than f…
FAST: Efficient Action Tokenization for Vision-Language-Action Models
Karl Pertsch, Kyle Stachowicz, Brian Ichter +6
Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behav…
FastRLAP: A System for Learning High-Speed Driving via Deep RL and Autonomous Practicing
Kyle Stachowicz, Dhruv Shah, Arjun Bhorkar +2
We present a system that enables an autonomous small-scale RC car to drive aggressively from visual observations using reinforcement learning (RL). Our system, FastRLAP (faster lap…