1 paper
Pratheswaran Hariharan, Haiping Xu, Donghui Yan
Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language response generation. However, these systems can…