5 citations · 5 across the 10 of their papers we have counts for
1 paper · 1 filter
Lujun Li, Lama Sleem, Niccolo Gentile +4
Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored. A natural extension of ``How m…