1 paper
Kiet T. Nguyen, Hanbo Shim, Jinwoo Kim +1
Vision-language models (VLMs) have achieved strong image and video understanding, yet their visual-spatial representations remain geometrically fragile, leading to failures in spat…