2 citations · 3 across the 5 of their papers we have counts for
7 papers · 1 filter
AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge
Noriaki Hirose, Catherine Glossop, Dhruv Shah +1
Robotic foundation models achieve strong generalization by leveraging internet-scale vision-language representations, but their massive computational cost creates a fundamental bot…
SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
Tian Gao, Celine Tan, Catherine Glossop +8
A fundamental challenge in autonomous driving is the integration of high-level, semantic reasoning for long-tail events with low-level, reactive control for robust driving. While l…
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
William Chen, Jagdeep Singh Bhatia, Catherine Glossop +6
Pretrained vision-language models (VLMs) can make semantic and visual inferences across diverse settings, providing valuable common-sense priors for robotic control. However, effec…
OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
Noriaki Hirose, Catherine Glossop, Dhruv Shah +1
Humans can flexibly interpret and compose different goal specifications, such as language instructions, spatial coordinates, or visual references, when navigating to a destination.…
Learning to Drive Anywhere with Model-Based Reannotation
Noriaki Hirose, Lydia Ignatova, Kyle Stachowicz +3
Developing broadly generalizable visual navigation policies for robots is a significant challenge, primarily constrained by the availability of large-scale, diverse training data.…
LeLaN: Learning A Language-Conditioned Navigation Policy from In-the-Wild Videos
Noriaki Hirose, Catherine Glossop, Ajay Sridhar +3
The world is filled with a wide variety of objects. For robots to be useful, they need the ability to find arbitrary objects described by people. In this paper, we present LeLaN(Le…