collaborators
Showing cs.ROShow all

7 papers · 1 filter

cs.RO2026

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

Yixiang Chen, Jiabing Yang, Yuan Xu +10

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture p…

cs.RO2026

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu +13

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for co…

cs.RO2026

WAM-Nav: Asymmetric Latent World-Action Modeling for Unified Visual Navigation

Ning Yang, Yan Huang, Kaiwen Peng +9

Visual navigation requires generating smooth and collision-free trajectories under complex geometric and physical constraints. Existing reactive policies that directly map observat…

cs.RO2026

SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Peiyan Li, Yixiang Chen, Yuan Xu +13

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. The…

cs.RO2026

FloorPlan-VLN: A New Paradigm for Floor Plan Guided Vision-Language Navigation

Kehan Chen, Yan Huang, Dong An +5

Existing Vision-Language Navigation (VLN) task requires agents to follow verbose instructions, ignoring some potentially useful global spatial priors, limiting their capability to…

cs.RO2026

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

Yixiang Chen, Peiyan Li, Jiabing Yang +8

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual an…