6 papers
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
Teqiang Zou, Hongliang Zeng, Yuxuan Nong +6
Most Vision-Language-Action (VLA) systems integrate a Vision-Language Model (VLM) for semantic reasoning with an action expert generating continuous action signals, yet both typica…
MLM: Learning Multi-task Loco-Manipulation Whole-Body Control for Quadruped Robot with Arm
Xin Liu, Bida Ma, Chenkun Qi +14
Whole-body loco-manipulation for quadruped robots with arms remains a challenging problem, particularly in achieving multi-task control. To address this, we propose MLM, a reinforc…
FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset
Kehui Liu, Zhongjie Jia, Yang Li +14
Data-driven robotic manipulation learning depends on large-scale, high-quality expert demonstration datasets. However, existing datasets, which primarily rely on human teleoperated…
COHERENT: Collaboration of Heterogeneous Multi-Robot System with Large Language Models
Kehui Liu, Zixin Tang, Dong Wang +3
Leveraging the powerful reasoning capabilities of large language models (LLMs), recent LLM-based robot task planning methods yield promising results. However, they mainly focus on…
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
Pingrui Zhang, Xianqiang Gao, Yuhan Wu +6
In mobile manipulation, navigation and manipulation are often treated as separate problems, resulting in a significant gap between merely approaching an object and engaging with it…
FastUMI: A Scalable and Hardware-Independent Universal Manipulation Interface with Dataset
Zhaxizhuoma, Kehui Liu, Chuyue Guan +15
Real-world manipulation data involving robotic arms is crucial for developing generalist action policies, yet such data remains scarce since existing data collection methods are hi…