10 papers · 1 filter
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
Li Wenjie, Yash Jangir, Ignacy Stepka +3
Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are ty…
IndoorR2X: Indoor Robot-to-Everything Coordination with LLM-Driven Planning
Fan Yang, Soumya Teotia, Shaunak A. Mehta +8
Although robot-to-robot (R2R) communication improves indoor scene understanding beyond what a single robot can achieve, R2R alone cannot overcome partial observability without subs…
ELASTIC: Efficiently Learning to Adaptively Scale Test-Time Compute for Generative Control Policies
Andrew Zou Li, Gokul Swamy, Yonatan Bisk +1
Generative control policies (GCPs), such as diffusion policies and flow-based vision-language-action models, enable test-time scaling in robot control. Test-time compute can be all…
Flatness Preserves Instruction Following in Vision-Language-Action Models
Haochen Zhang, Yonatan Bisk
Vision-language-action (VLA) models have the potential for open-world generalization by leveraging pretrained vision-language representations, yet downstream finetuning on limited…
Inductive Generalization for Robotic Manipulation
Annabella Macaluso, Haochen Zhang, Ishaan Masilamony +2
Understanding the generalization capabilities of visuomotor policies is essential in the development of capable robotic agents. Generalizable models learn structures that transfer…
RobotArena : Scalable Robot Benchmarking via Real-to-Sim Translation
Yash Jangir, Yidi Zhang, Pang-Chi Lo +7
The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot…