9 papers
StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
Xiao Liu, Yuguang Yang, Xi Wang +6
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the fut…
4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields
Lishan Yang, Wenxuan Song, Xi Wang +14
Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel s…
EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang +3
Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured rep…
MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight
Zehua Fan, Junjie He, Wenxuan Song +14
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands sim…
Modeling Subjective Urban Perception with Human Gaze
Lin Che, Xi Wang, Marc Pollefeys +3
Urban perception describes how people subjectively evaluate urban environments, shaping how cities are experienced and understood. Existing computational approaches primarily model…
PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation
Zehua Fan, Wenqi Lyu, Wenxuan Song +12
Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but als…