activity
20242026
collaborators

5 papers

cs.RO2026

InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26

Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…

cs.CV2026

EgoSim: Egocentric World Simulator for Embodied Interaction Generation

Jinkun Hao, Mingda Jia, Ruiyan Wang +7

We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updates the underlying 3D scene state for cont…

cs.RO2025

Wavelet Policy: Imitation Learning in the Scale Domain with World Prior Memory

Changchuan Yang, Haoxuan Xu, Yuhang Dong +3

Conventional visuomotor imitation learning usually predicts future robot actions directly in the time domain. Such formulations often have limited physical scene awareness and weak…

cs.AI2025

ImitDiff: Transferring Foundation-Model Priors for Distraction Robust Visuomotor Policy

Yuhang Dong, Haizhou Ge, Yupei Zeng +9

Visuomotor imitation learning policies enable robots to efficiently acquire manipulation skills from visual demonstrations. However, as scene complexity and visual distractions inc…

cs.LG2024

Bridging the Resource Gap: Deploying Advanced Imitation Learning Models onto Affordable Embedded Platforms

Haizhou Ge, Ruixiang Wang, Zhu-ang Xu +7

Advanced imitation learning with structures like the transformer is increasingly demonstrating its advantages in robotics. However, deploying these large-scale models on embedded p…