10 papers · 1 filter
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
Xin Wu, Zhixuan Liang, Yue Ma +3
Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable cha…
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Yuhao Zhang, Wanxi Dong, Yue Shi +13
Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact-rich actions. While large-scale 3D vision models provide…
UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data
Sizhe Yang, Yiman Xie, Zhixuan Liang +4
Grasping is a fundamental capability for robots to interact with the physical world. Humans, equipped with two hands, autonomously select appropriate grasp strategies based on the…
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
Weijie Shen, Yitian Liu, Yuhao Wu +10
Vision-Language-Action (VLA) models are experiencing rapid development and demonstrating promising capabilities in robotic manipulation tasks. However, scaling up VLA models presen…
HyCodePolicy: Hybrid Language Controllers for Multimodal Monitoring and Decision in Embodied Agents
Yibin Liu, Zhixuan Liang, Zanxin Chen +7
Recent advances in multimodal large language models (MLLMs) have enabled richer perceptual grounding for code policy generation in embodied agents. However, most existing systems l…
Benchmarking Generalizable Bimanual Manipulation: RoboTwin Dual-Arm Collaboration Challenge at CVPR 2025 MEIS Workshop
Tianxing Chen, Kaixuan Wang, Zhaohui Yang +96
Embodied Artificial Intelligence (Embodied AI) is an emerging frontier in robotics, driven by the need for autonomous systems that can perceive, reason, and act in complex physical…