1 paper
Xiaoda Yang, Shuai Yang, Can Wang +9
Vision-Language Models (VLMs) have made significant strides in static image understanding but continue to face critical hurdles in spatiotemporal reasoning. A major bottleneck is "…