7 papers
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
Qifeng Zhang, Kaixiang Huang, Heng Dong +6
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awarenes…
AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
Jinxuan Zhu, Chenrui Tie, Xinyi Cao +8
Non-prehensile (NP) manipulation, in which robots alter object states without forming stable grasps (for example, pushing, poking, or sliding), significantly broadens robotic manip…
Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals via Vision-Language Models
Chenrui Tie, Shengxiang Sun, Yudi Lin +9
Assembly hinges on reliably forming connections between parts; yet most robotic approaches plan assembly sequences and part poses while treating connectors as an afterthought. Conn…
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
Zixuan Chen, Nga Teng Chan, Yiwen Hou +8
Bimanual manipulation is a fundamental robotic skill that requires continuous and precise coordination between two arms. While imitation learning (IL) is the dominant paradigm for…
LISN: Language-Instructed Social Navigation with VLM-based Controller Modulating
Junting Chen, Yunchuan Li, Panfeng Jiang +5
Towards human-robot coexistence, socially aware navigation is significant for mobile robots. Yet existing studies on this area focus mainly on path efficiency and pedestrian collis…
Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models
Chenrui Tie, Shengxiang Sun, Jinxuan Zhu +7
Humans possess an extraordinary ability to understand and execute complex manipulation tasks by interpreting abstract instruction manuals. For robots, however, this capability rema…