1 paper · 1 filter
Artemis Panagopoulou, Aveek Purohit, Achin Kulshrestha +2
While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, suc…