1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Yiming Qin, Bomin Wei, Jiaxin Ge +4
Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual perception, e.g., spatial reasoning and g…