activity
20242026
collaborators

9 papers

cs.RO2026

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

Catherine Glossop, William Chen, Arjun Bhorkar +2

Generalist robots should be able to understand and follow user instructions. Despite providing a powerful architecture for mapping open-vocabulary language instructions to robot ac…

cs.LG2026

: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities

Physical Intelligence, Bo Ai, Ali Amin +85

We present a new robotic foundation model, called , that can enable strong out-of-the-box performance in a wide range of scenarios. can follow diverse language…

cs.RO2026

Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control

William Chen, Jagdeep Singh Bhatia, Catherine Glossop +6

Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effec…

cs.RO2026

AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge

Noriaki Hirose, Catherine Glossop, Dhruv Shah +1

Robotic foundation models achieve strong generalization by leveraging internet-scale vision-language representations, but their massive computational cost creates a fundamental bot…

cs.RO2026

SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios

Tian Gao, Celine Tan, Catherine Glossop +8

A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While l…

cs.RO2025

Learning to Drive Anywhere with Model-Based Reannotation

Noriaki Hirose, Lydia Ignatova, Kyle Stachowicz +3

Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data.…