1 paper · 1 filter
Chen Yang, Shenxiang Zeng, Haoyang Zhao +6
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-language models (VLMs) often struggle to int…