25 citations · 25 across the 1 of their papers we have counts for
1 paper
Yucheng Han, Chi Zhang, Xin Chen +5
Multi-modal large language models have demonstrated impressive performances on most vision-language tasks. However, the model generally lacks the understanding capabilities for spe…