1 paper
Zeju Li, Chao Zhang, Xiaoyan Wang +4
The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3…