3 papers
cs.RO2026
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
Jingkai Wang, Zihan Tang, Gu Zhang +7
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this c…
cs.RO2026
UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
Gu Zhang, Qicheng Xu, Haozhe Zhang +16
Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of contro…
cs.CV2026
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
Maanping Shao, Feihong Zhang, Gu Zhang +3
Visuomotor policies often leverage large pre-trained Vision Transformers (ViTs) for their powerful generalization capabilities. However, their significant data requirements present…