15 papers
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
Li Wenjie, Yash Jangir, Ignacy Stepka +3
Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are ty…
IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning
Fan Yang, Soumya Teotia, Shaunak A. Mehta +8
Although robot-to-robot (R2R) communication improves indoor scene understanding beyond what a single robot can achieve, R2R alone cannot overcome partial observability without subs…
ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies
Andrew Zou Li, Gokul Swamy, Yonatan Bisk +1
Generative control policies (GCPs), such as diffusion policies and flow-based vision-language-action models, enable test-time scaling in robot control. Test-time compute can be all…
Flatness Preserves Instruction Following in Vision-Language-Action Models
Haochen Zhang, Yonatan Bisk
Vision-language-action (VLA) models have the potential for open-world generalization by leveraging pretrained vision-language representations, yet downstream finetuning on limited…
Inductive Generalization for Robotic Manipulation
Annabella Macaluso, Haochen Zhang, Ishaan Masilamony +2
Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer…
Evaluation of ML Resource Utilization Requires Model Life Cycle Assessment
Jared Fernandez, Clara Na, Yonatan Bisk +2
Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers, policy makers, and users to…