8 papers
Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models
Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian +4
Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent…
Do World Action Models Generalize Better than VLAs? A Robustness Study
Zhanguang Zhang, Zhiyuan Li, Behnam Rahmati +11
Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response…
How VLAs (Really) Work In Open-World Environments
Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4
Vision-language-action models (VLAs) have been extensively used in robotics applications, achieving great success in various manipulation problems. More recently, VLAs have been us…
Distracted Robot: How Visual Clutter Undermine Robotic Manipulation
Amir Rasouli, Montgomery Alban, Sajjad Pakdamansavoji +4
In this work, we propose an evaluation protocol for examining the performance of robotic manipulation policies in cluttered scenes. Contrary to prior works, we approach evaluation…
Improving Robotic Manipulation Robustness via NICE Scene Surgery
Sajjad Pakdamansavoji, Mozhgan Pourkeshavarz, Adam Sigal +3
Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety.…
CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision Avoidance
Rui Heng Yang, Xuan Zhao, Leo Maxime Brunswic +5
In robotics, diffusion models can capture multi-modal trajectories from demonstrations, making them a transformative approach in imitation learning. However, achieving optimal perf…