From the 1 of 18 linked papers with an AI index.
17 papers
KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation
Yaping Li, Zhaxizhuoma, Qiaojun Yu +3
Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-A…
Scaling Behavior Foundation Model for Humanoid Robots
Weishuai Zeng, Kangning Yin, Xiaojie Niu +15
The paper proposes a scalable behavior foundation model for humanoid robots that uses a motion‑tracking learning paradigm, coordinated on‑policy rollouts and diverse reference moti…
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Sihan Yang, Runsen Xu, Yiman Xie +10
Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relati…
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
Runsen Xu, Weiyao Wang, Hao Tang +5
Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physi…
Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction
Sizhe Yang, Linning Xu, Hao Li +4
3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise an…
FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model
Xiaoxu Xu, Hao Li, Jinhui Ye +7
Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geomet…