1 paper
Seokju Cho, Abhishek Badki, Hang Su +5
Despite recent advances, Vision Language Models (VLMs) still struggle to grasp the dynamics of the world. We note that the ability to reason about a 4D scene, challenging in itself…