1 paper
Ge Yan, Jinghao Liu, Yuzhi Fan +4
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for t…