1 paper
Haoran Zhou, Zihan Zhang, Hao Chen
Multimodal Large Language Models (MLLMs) have made significant strides by combining visual recognition and language understanding to generate content that is both coherent and cont…