1 paper
Dianyu Wang, Peirong Zhang, Xuyang Li +2
Recent multimodal large language models (MLLMs) have shown strong cross-modal understanding and coordinate generation abilities in visual grounding. However, transferring these abi…