1 paper · 1 filter
Xiaoshuang Huang, Haifeng Huang, Lingdong Shen +4
With the rapid development of multimodal large language models (MLLMs), especially their capabilities in visual chat through refer and ground functionalities, their significance is…