23 papers
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ye Wang, Pei Lin, Xiong-Hui Chen +12
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, an…
Human-Centric Transferable Tactile Pre-Training for Dexterous Robotic Manipulation
Chi Zhang, Penglin Cai, Ziheng Xi +6
As an essential modality for dexterous and contact-rich tasks, tactile sensing provides precise force feedback that cannot be reliably inferred from vision. However, limited by har…
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
Wanpeng Zhang, Ye Wang, Hao Luo +6
Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an internal signal for judging whether a given prediction is reliable. Distributio…
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Jiazhao Zhang, Gengze Zhou, Hale Yin +32
Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search…
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
Boyu Li, Chaoyi Xu, Haoqi Yuan +5
Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse…
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Haoqi Yuan, Zhixuan Liang, Anzhe Chen +20
Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we i…