3 papers
cs.RO2026
DreamWAM: Beyond RGB Future Prediction for World Action Models
Shanglin Yuan, Weiheng Zhao, Xin Shi +6
World Action Models (WAMs) learn action-relevant representations by predicting how the observed world will evolve. Most existing WAMs define this future in RGB space, where task-re…
cs.RO2026
MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model
Shanglin Yuan, Weiheng Zhao, Xianda Guo +4
Vision-language-action (VLA) models increasingly condition robot policies on history, depth, or 4D features to resolve ambiguity in long-horizon manipulation. However, more spatiot…
cs.CV2025
MobileI2V: Fast and High-Resolution Image-to-Video on Mobile Devices
Shuai Zhang, Bao Tang, Siyuan Yu +7
Recently, video generation has witnessed rapid advancements, drawing increasing attention to image-to-video (I2V) synthesis on mobile devices. However, the substantial computationa…