works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.RO2026

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Yixiang Chen, Jiabing Yang, Yuan Xu +10

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture p…

cs.RO2026

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu +13

The paper introduces FlowWAM, a dual‑stream diffusion model that uses optical flow as a unified video‑native representation of actions, enabling both action prediction and world mo…

cs.RO2026

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Ning Yang, Yan Huang, Kaiwen Peng +9

Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policies that directly map observat…

cs.CV2026

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Shiqiang Lang, Jing Liu, Haoyang He +6

Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon tasks such as autonomous driving…

cs.RO2026

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Peiyan Li, Yixiang Chen, Yuan Xu +13

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. The…

cs.RO2026

EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation

Yuan Xu, Jiabing Yang, Xiaofeng Wang +16

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-…