7 papers
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Zihan Guo, Xiaoqi Li
The paper proposes a cloud‑edge framework that combines an enhanced YOLO‑based gesture detector with coordinated large language model and vision‑language model agents to enable rob…
ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
Chenyang Gu, Jiaming Liu, Hao Chen +9
Vision-Language-Action (VLA) models have recently emerged, demonstrating strong generalization in robotic scene understanding and manipulation. However, when confronted with long-h…
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
Rongyu Zhang, Jiaming Liu, Xiaoqi Li +5
Vision-centric Bird's Eye View (BEV) perception holds considerable promise for autonomous driving. Recent studies have prioritized efficiency or accuracy enhancements, yet the issu…
RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot
Liang Heng, Xiaoqi Li, Shangqing Mao +9
Recent advancements in imitation learning have shown promising results in robotic manipulation, driven by the availability of high-quality training data. To improve data collection…
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
Jiaming Liu, Hao Chen, Pengju An +12
A fundamental objective of manipulation policy design is to endow robots to comprehend human instructions, reason about scene cues, and execute generalized actions in dynamic envir…
Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
Hao Chen, Jiaming Liu, Chenyang Gu +8
Generalized policy and execution efficiency constitute the two critical challenges in robotic manipulation. While recent foundation policies benefit from the common-sense reasoning…