4 papers
Flatness Preserves Instruction Following in Vision-Language-Action Models
Haochen Zhang, Yonatan Bisk
Vision-language-action (VLA) models have the potential for open-world generalization by leveraging pretrained vision-language representations, yet downstream finetuning on limited…
Inductive Generalization for Robotic Manipulation
Annabella Macaluso, Haochen Zhang, Ishaan Masilamony +2
Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer…
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
Nader Zantout, Haochen Zhang, Pujith Kachana +4
Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is…
IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes
Haochen Zhang, Nader Zantout, Pujith Kachana +2
With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can…