Showing cs.ROShow all
2 papers · 1 filter
cs.RO2026
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
Chenhao Qiu, Ruixiang Wang, Runyi Zhao +7
World Action Models (WAMs) improve robot policy learning by jointly modeling future visual dynamics and actions. However, their scalability and generalization remain constrained by…
cs.RO2026
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
Songen Gu, Yunuo Cai, Tianyu Wang +2
Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors…