1 paper · 1 filter
Zhoutong Ye, Mingze Sun, Huan-ang Gao +9
Large multimodal models (LMMs) have demonstrated significant potential as generalists in vision-language (VL) tasks. However, adoption of LMMs in real-world tasks is hindered by th…