1 paper · 1 filter
Shezheng Song, Chengxiang He, Shan Zhao +4
Multimodal large language models (MLLMs) have shown remarkable progress in high-level semantic tasks such as visual question answering, image captioning, and emotion recognition. H…