1 paper
Yuchen Li, Amanmeet Garg, Shalini Chaudhuri +2
Large Vision Language Models (LVLMs) excel at semantic understanding but struggle with fine grained spatial grounding, as the model must implicitly infer complex geometry without e…