1 paper
Xinkai Wang, Beibei Li, Zerui Shao +3
Multimodal large language models (MLLMs) have become integral to a wide range of real-world applications by jointly reasoning over text and visual inputs. However, despite recent a…