17 citations · 35 across the 24 of their papers we have counts for
1 paper · 1 filter
Yifu Qiu, Yftah Ziser, Anna Korhonen +2
Can unified vision-language models (VLMs) perform forward dynamics prediction (FDP), i.e., predicting the future state (in image form) given the previous observation and an action…