2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Bin Wang, Fan Wu, Xiao Han +8
The integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality in…