10 papers
A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation
Yi-Xiang He, Lan Wei, Haoming Cen +6
Multi-robot systems provide the parallelism and redundancy necessary for long-horizon tasks, while Large Language Models (LLMs) offer the reasoning capabilities to decompose these…
HATS: A Human-Agent Teleoperation System for Multi-Arm Data Collection
Zesen Lin, Jian-Jian Jiang, Haoming Cen +3
Many real-world manipulation scenarios, such as handling complex collaborative tasks and dealing with large workspaces, require coordination of more than two robotic arms. Conseque…
Task Editing for Generalizable 3D Visuomotor Policy Learning
Jian-Jian Jiang, YiHan Yang, Lan Wei +6
3D visuomotor policies offer a promising direction for complex robotic manipulation, as depth maps and point clouds provide rich geometric information for spatial reasoning. Howeve…
VLANeXt: Recipes for Building Strong VLA Models
Xiao-Ming Wu, Bin Fan, Kang Liao +6
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for gen…
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
Zhuohao Li, Yinghao Li, Jian-Jian Jiang +6
Vision-Language-Action (VLA) models have advanced robotic manipulation by combining vision, language, and proprioception to predict actions. However, previous methods fuse proprioc…
ProEdit: Inversion-based Editing From Prompts Done Right
Zhi Ouyang, Dian Zheng, Xiao-Ming Wu +4
Inversion-based visual editing provides an effective and training-free way to edit an image or a video based on user instructions. Existing methods typically inject source image in…