1 paper
Niyati Rawal, Sushant Ravva, Shah Alam Abir +5
Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spati…