2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Yi Lu, Jiawang Cao, Yongliang Wu +6
Multi-modal Large Language Models (MLLMs) have demonstrated remarkable reasoning capability while lack explicit mechanisms for visual grounding and segmentation, creating a gap bet…