collaborators

8 papers

cs.CV2026

Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models

Hunter Schofield, Mohammed Elmahgiubi, Mohammad Mahdavian +4

Spatial understanding is fundamental to embodied intelligence, underpinning applications such as robotic manipulation, embodied navigation, and autonomous driving. Although recent…

cs.RO2026

Do World Action Models Generalize Better than VLAs? A Robustness Study

Zhanguang Zhang, Zhiyuan Li, Behnam Rahmati +11

Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predicting how it will evolve in response…

cs.RO2026

How VLAs (Really) Work In Open-World Environments

Amir Rasouli, Yangzheng Wu, Zhiyuan Li +4

Vision-language-action models (VLAs) have been extensively used in robotics applications, achieving great success in various manipulation problems. More recently, VLAs have been us…

cs.RO2025

Distracted Robot: How Visual Clutter Undermine Robotic Manipulation

Amir Rasouli, Montgomery Alban, Sajjad Pakdamansavoji +4

In this work, we propose an evaluation protocol for examining the performance of robotic manipulation policies in cluttered scenes. Contrary to prior works, we approach evaluation…

cs.RO2025

Improving Robotic Manipulation Robustness via NICE Scene Surgery

Sajjad Pakdamansavoji, Mozhgan Pourkeshavarz, Adam Sigal +3

Learning robust visuomotor policies for robotic manipulation remains a challenge in real-world settings, where visual distractors can significantly degrade performance and safety.…

cs.RO2025

CAPE: Context-Aware Diffusion Policy Via Proximal Mode Expansion for Collision Avoidance

Rui Heng Yang, Xuan Zhao, Leo Maxime Brunswic +5

In robotics, diffusion models can capture multi-modal trajectories from demonstrations, making them a transformative approach in imitation learning. However, achieving optimal perf…