12 papers
RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation
Ziqin Wang, Hao Li, Weijun Wang +5
Existing robot datasets remain expensive to curate, embodiment-specific, and insufficiently annotated with the fine-grained structure required for generalizable reasoning, executio…
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
Haoxiang Ma, Junhao Cai, Xiaoxu Xu +26
Unified models for robot manipulation aim to equip one policy with both the semantic priors of pretrained VLMs and the physical dynamics learned through future prediction. In pract…
EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
Ning Gao, Jinliang Zheng, Xing Gao +22
We present EBench, a simulation benchmark that diagnoses generalist mobile manipulation policies beyond a single success-rate scalar. EBench comprises 26 diverse and challenging ma…
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
Sizhe Yang, Linning Xu, Hao Li +4
3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise an…
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
Xiaoxu Xu, Hao Li, Jinhui Ye +7
Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geomet…
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
Shuai Yang, Hao Li, Bin Wang +7
To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often s…