3 papers
cs.RO2026
Co-VLA: Coordination-Aware Structured Action Modeling for Dual-Arm Vision-Language-Action Systems
Yandong Wang, Jiaqian Yu, Xiongfeng Peng +8
Vision-language-action (VLA) models show strong capabilities in single and dual-arm robotic manipulation. Prior works show coordinated bimanual behaviors can emerge from end-to-end…
cs.AI2025
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
Lu Xu, Jiaqian Yu, Xiongfeng Peng +7
To meet the growing demand for smarter, faster, and more efficient embodied AI solutions, we introduce a novel Mixture-of-Expert (MoE) method that significantly boosts reasoning an…
cs.CV2025
How Can Objects Help Video-Language Understanding?
Zitian Tang, Shijie Wang, Junho Cho +2
Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which obj…