1 paper
Xingrui Wang, Wufei Ma, Angtian Wang +3
For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-lev…