2 papers
cs.RO2026
RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models
Hao Wu, Yuqi Li, Yuan Gao +10
Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities t…
cs.CV2025
PastNet: Introducing Physical Inductive Biases for Spatio-temporal Video Prediction
Hao Wu, Fan Xu, Chong Chen +3
In this paper, we investigate the challenge of spatio-temporal video prediction task, which involves generating future video frames based on historical spatio-temporal observation…