works on

From the 1 of 18 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.RO2026

KAI: A Kinematic-Aware Interface for Data-Efficient Articulated Object Manipulation

Yaping Li, Zhaxizhuoma, Qiaojun Yu +3

Articulated object manipulation requires an understanding of kinematic structure that is difficult and costly to learn from robot demonstrations alone. We introduce the Kinematic-A…

cs.RO2026

Scaling Behavior Foundation Model for Humanoid Robots

Weishuai Zeng, Kangning Yin, Xiaojie Niu +15

The paper proposes a scalable behavior foundation model for humanoid robots that uses a motion‑tracking learning paradigm, coordinated on‑policy rollouts and diverse reference moti…

cs.CV2026

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Sihan Yang, Runsen Xu, Yiman Xie +10

Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relati…

cs.CV2026

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

Runsen Xu, Weiyao Wang, Hao Tang +5

Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single images, leaving them ill-suited for physi…

cs.RO2026

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

Sizhe Yang, Linning Xu, Hao Li +4

3D spatial perception is fundamental to generalizable robotic manipulation, yet obtaining reliable, high-quality 3D geometry remains challenging. Depth sensors suffer from noise an…

cs.RO2026

FutureVLA: Joint Visuomotor Prediction for Vision-Language-Action Model

Xiaoxu Xu, Hao Li, Jinhui Ye +7

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geomet…