works on

From the 1 of 10 linked papers with an AI index.

collaborators
Showing cs.ROShow all

7 papers · 1 filter

cs.RO2026

XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations

Shichao Fan, Kun Wu, Zhengping Che +12

Recent progress in large-scale robotic datasets and vision-language models (VLMs) has advanced research on vision-language-action (VLA) models. However, existing VLA models still f…

cs.RO2026

UMI-Bench 1.0: An Open and Reproducible Real-World Benchmark for Tabletop Robotic Manipulation with UMI Data

Shi Jin, Yuntian Wang, Yuhui Duan +16

Real-robot evaluation is essential for understanding whether learned manipulation policies can operate reliably outside curated demonstrations. This need is particularly pressing f…

cs.RO2026

Generalizable and Actionable Parts Pose Estimation with Symmetry Annotation-Free Learning Strategy

Wenxiao Chen, Xueyu Yuan, Liu Liu +2

Urgently needed generalizable robot object interaction and manipulation requires high-quality Cross-Category object perception. As a pioneer of this area, Generalizable and Actiona…

cs.RO2026

EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations

Justin Yu, Yide Shentu, Di Wu +3

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodim…

cs.RO2026

RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipulation Dataset for Generalizable Embodied Intelligence

Chengkai Hou, Kun Wu, Jiaming Liu +30

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstration…

cs.RO2026

RoboGene: Boosting VLA Pre-training via Diversity-Driven Agentic Framework for Real-World Task Generation

Yixue Zhang, Kun Wu, Zhi Gao +12

The pursuit of general-purpose robotic manipulation is hindered by the scarcity of diverse, real-world interaction data. Unlike data collection from web in vision or language, robo…