4 citations · 4 across the 1 of their papers we have counts for
1 paper
Leyang Shen, Gongwei Chen, Rui Shao +2
Multimodal large language models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. However, a generalist MLLM typically underperforms compared…