1 paper
Fanxiao Li, Jiaying Wu, Canyuan He +1
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in visual reasoning and text generation. While previous studies have explored the application of…