papers

Publications (17)

cs.RO2023

ViNT: A Foundation Model for Visual Navigation

Dhruv Shah, Ajay Sridhar, Nitish Dashora +4

General-purpose pre-trained models ("foundation models") have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that a…

cs.LG2026

: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Physical Intelligence, Bo Ai, Ali Amin +85

We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language…

cs.RO2024

RACER: Epistemic Risk-Sensitive RL Enables Fast Driving with Fewer Crashes

Kyle Stachowicz, Sergey Levine

Reinforcement learning provides an appealing framework for robotic control due to its ability to learn expressive policies purely through real-world interaction. However, this requ…

cs.RO2025

Learning to Drive Anywhere with Model-Based Reannotation

Noriaki Hirose, Lydia Ignatova, Kyle Stachowicz +3

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data.…

cs.RO2025

Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding

Joshua Jones, Oier Mees, Carmelo Sferrazza +3

Interacting with the world is a multi-sensory experience: achieving effective general-purpose interaction requires making use of all available modalities -- including vision, touch…

math.OC2022

Multimodal Maximum Entropy Dynamic Games

Oswin So, Kyle Stachowicz, Evangelos A. Theodorou

Environments with multi-agent interactions often result a rich set of modalities of behavior between agents due to the inherent suboptimality of decision making processes when agen…

cs.RO2026

SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios

Tian Gao, Celine Tan, Catherine Glossop +8

A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While l…

cs.RO2024

SELFI: Autonomous Self-Improvement with Reinforcement Learning for Social Navigation

Noriaki Hirose, Dhruv Shah, Kyle Stachowicz +2

Autonomous self-improving robots that interact and improve with experience are key to the real-world deployment of robotic systems. In this paper, we propose an online learning met…

cs.RO2026

MEM: Multi-Scale Embodied Memory for Vision Language Action Models

Marcel Torne, Karl Pertsch, Homer Walke +14

Conventionally, memory in end-to-end robotic learning involves inputting a sequence of past observations into the learned policy. However, in complex multi-stage real-world tasks,…

cs.RO2024

Traversability-Aware Legged Navigation by Learning from Real-World Visual Data

Hongbo Zhang, Zhongyu Li, Xuanqi Zeng +9

The enhanced mobility brought by legged locomotion empowers quadrupedal robots to navigate through complex and unstructured environments. However, optimizing agile locomotion while…

cs.RO2022

Safety Embedded Differential Dynamic Programming Using Discrete Barrier States

Hassan Almubarak, Kyle Stachowicz, Nader Sadegh +1

Certified safe control is a growing challenge in robotics, especially when performance and safety objectives must be concurrently achieved. In this work, we extend the barrier stat…

cs.LG2025

: a Vision-Language-Action Model with Open-World Generalization

Physical Intelligence, Kevin Black, Noah Brown +33

In order for robots to be useful, they must perform practically relevant tasks in the real world, outside of the lab. While vision-language-action (VLA) models have demonstrated im…

cs.LG2025

: a VLA That Learns From Experience

Physical Intelligence, Ali Amin, Raichelle Aniceto +53

We study how vision-language-action (VLA) models can improve through real-world deployments via reinforcement learning (RL). We present a general-purpose method, RL with Experience…

cs.RO2026

Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?

Filippo Lazzati, Kyle Stachowicz, William Chen +3

Action chunking---predicting and executing multiple actions instead of a single action---has proven to be a critical component for learning effective robotic control policies. Howe…

cs.RO2021

Optimal-Horizon Model-Predictive Control with Differential Dynamic Programming

Kyle Stachowicz, Evangelos A. Theodorou

We present an algorithm, based on the Differential Dynamic Programming framework, to handle trajectory optimization problems in which the horizon is determined online rather than f…

cs.RO2025

FAST: Efficient Action Tokenization for Vision-Language-Action Models

Karl Pertsch, Kyle Stachowicz, Brian Ichter +6

Autoregressive sequence models, such as Transformer-based vision-language action (VLA) policies, can be tremendously effective for capturing complex and generalizable robotic behav…

cs.RO2023

FastRLAP: A System for Learning High-Speed Driving via Deep RL and Autonomous Practicing

Kyle Stachowicz, Dhruv Shah, Arjun Bhorkar +2

We present a system that enables an autonomous small-scale RC car to drive aggressively from visual observations using reinforcement learning (RL). Our system, FastRLAP (faster lap…