Showing cs.ROShow all
3 papers · 1 filter
cs.RO2026
Flatness Preserves Instruction Following in Vision-Language-Action Models
Haochen Zhang, Yonatan Bisk
Vision-language-action (VLA) models have the potential for open-world generalization by leveraging pretrained vision-language representations, yet downstream finetuning on limited…
cs.RO2026
Inductive Generalization for Robotic Manipulation
Annabella Macaluso, Haochen Zhang, Ishaan Masilamony +2
Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer…
cs.RO2024
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
Haochen Zhang, Nader Zantout, Pujith Kachana +3
With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embo…