12 papers
Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
Ye Wang, Pei Lin, Xiong-Hui Chen +12
Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, an…
HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration
Chang Liu, Jiawei Zhang, Tao Zhang +3
Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-wor…
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models
Wanpeng Zhang, Ye Wang, Hao Luo +6
Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an internal signal for judging whether a given prediction is reliable. Distributio…
GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments
Jiawei Zhang, Yiming Yan, Chao Liang +8
Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environments offer a compelling alternative by enabling large-scale, cost-e…
EasyMimic: A Low-Cost Framework for Robot Imitation Learning from Human Videos
Tao Zhang, Song Xia, Ye Wang +1
Robot imitation learning is often hindered by the high cost of collecting large-scale, real-world data. This challenge is especially significant for low-cost robots designed for ho…
Rethinking Visual-Language-Action Model Scaling: Alignment, Mixture, and Regularization
Ye Wang, Sipeng Zheng, Hao Luo +9
While Vision-Language-Action (VLA) models show strong promise for generalist robot control, it remains unclear whether -- and under what conditions -- the standard "scale data" rec…