activity
20242026
collaborators

8 papers

cs.RO2026

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Xiao Liu, Yuguang Yang, Xi Wang +6

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…

cs.RO2026

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

Lishan Yang, Wenxuan Song, Xi Wang +14

Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…

cs.CV2026

EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang +3

Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured rep…

cs.CV2026

MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

Zehua Fan, Junjie He, Wenxuan Song +14

World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…

cs.CV2026

Modeling Subjective Urban Perception with Human Gaze

Lin Che, Xi Wang, Marc Pollefeys +3

Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model…

cs.CV2026

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

Zehua Fan, Wenqi Lyu, Wenxuan Song +12

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but als…