4 citations · 5 across the 7 of their papers we have counts for
1 paper · 2 filters
Fanrui Zhang, Jiawei Liu, Jiaying Zhu +4
Multimodal Large Language Models (MLLMs), such as GPT4o, have shown strong capabilities in visual reasoning and explanation generation. However, despite these strengths, they face…