7 citations · 7 across the 1 of their papers we have counts for
1 paper
Lai Wei, Xiaozhe Li, Zihao Jiang +2
Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using supervised vision-language instruction data…