1 paper
Guannan Lv, Ren Nie, Hongjian Dou +1
Multimodal Large Language Models (MLLMs) have increasingly localized and interleaved visual evidence for deliberative reasoning. Grounding-based approaches typically focus on regio…