2 citations · 2 across the 1 of their papers we have counts for
1 paper
Junxiao Xue, Quan Deng, Fei Yu +3
Multimodal large language models (MLLMs), such as GPT-4o, Gemini, LLaVA, and Flamingo, have made significant progress in integrating visual and textual modalities, excelling in tas…