1 paper
Yuanze Lin, Yunsheng Li, Dongdong Chen +4
In recent years, multimodal large language models (MLLMs) have made significant strides by training on vast high-quality image-text datasets, enabling them to generally understand…