1 citations · 2 across the 9 of their papers we have counts for
1 paper · 1 filter
Boyong Wu, Sanghwan Kim, Zeynep Akata
Multimodal Large Language Models (MLLMs) are increasingly applied to pixel-level vision tasks, yet their intrinsic capacity for spatial understanding remains poorly understood. We…