3 papers
cs.RO2026
G0.5: One Autoregressive Stream for Robot Reasoning and Action
Yicheng Liu, Zibin Dong, Baijun Ye +24
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder r…
cs.RO2026
TacUMI: A Multi-Modal Universal Manipulation Interface for Contact-Rich Tasks
Tailai Cheng, Kejia Chen, Lingyun Chen +8
Task decomposition is critical for understanding and learning complex long-horizon manipulation tasks. Especially for tasks involving rich physical interactions, relying solely on…
cs.RO2025
Multi-Robot Assembly of Deformable Linear Objects Using Multi-Modal Perception
Kejia Chen, Celina Dettmering, Florian Pachler +7
Industrial assembly of deformable linear objects (DLOs) such as cables offers great potential for many industries. However, DLOs pose several challenges for robot-based automation…