5 papers
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation
Jingkai Wang, Zihan Tang, Gu Zhang +7
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this c…
UniDex: A Robot Foundation Suite for Universal Dexterous Hand Control from Egocentric Human Videos
Gu Zhang, Qicheng Xu, Haozhe Zhang +16
Dexterous manipulation remains challenging due to the cost of collecting real-robot teleoperation data, the heterogeneity of hand embodiments, and the high dimensionality of contro…
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
Maanping Shao, Feihong Zhang, Gu Zhang +3
Visuomotor policies often leverage large pre-trained Vision Transformers (ViTs) for their powerful generalization capabilities. However, their significant data requirements present…
ArrayBot: Reinforcement Learning for Generalizable Distributed Manipulation through Touch
Zhengrong Xue, Han Zhang, Jingwen Cheng +5
We present ArrayBot, a distributed manipulation system consisting of a array of vertically sliding pillars integrated with tactile sensors, which can simultaneously…
Learning to Manipulate Anywhere: A Visual Generalizable Framework For Reinforcement Learning
Zhecheng Yuan, Tianming Wei, Shuiqi Cheng +3
Can we endow visuomotor robots with generalization capabilities to operate in diverse open-world scenarios? In this paper, we propose \textbf{Maniwhere}, a generalizable framework…