6 papers
SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
Yang Xu, Gurpreet Singh Mukker, Raymond Wang +3
Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to…
LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset
Royden Wagner, Omer Sahin Tas, Jaime Villa +18
In real-world domains such as self-driving, generalization to rare scenarios remains a fundamental challenge. To address this, we introduce a new dataset designed for end-to-end dr…
SAFE: Multitask Failure Detection for Vision-Language-Action Models
Qiao Gu, Yuanliang Ju, Shengxiang Sun +4
While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on no…
Box Pose and Shape Estimation and Domain Adaptation for Large-Scale Warehouse Automation
Xihang Yu, Rajat Talak, Jingnan Shi +3
Modern warehouse automation systems rely on fleets of intelligent robots that generate vast amounts of data -- most of which remains unannotated. This paper develops a self-supervi…
MORE: Mobile Manipulation Rearrangement Through Grounded Language Reasoning
Mohammad Mohammadi, Daniel Honerkamp, Martin Büchner +5
Autonomous long-horizon mobile manipulation encompasses a multitude of challenges, including scene dynamics, unexplored areas, and error recovery. Recent works have leveraged found…
GeoMatch++: Morphology Conditioned Geometry Matching for Multi-Embodiment Grasping
Yunze Wei, Maria Attarian, Igor Gilitschenski
Despite recent progress on multi-finger dexterous grasping, current methods focus on single grippers and unseen objects, and even the ones that explore cross-embodiment, often fail…