1 paper
Sirui Cheng, Siyu Zhang, Jiayi Wu +1
Within the multimodal field, large vision-language models (LVLMs) have made significant progress due to their strong perception and reasoning capabilities in the visual and languag…