2 citations · 2 across the 11 of their papers we have counts for
1 paper · 1 filter
Wenke Huang, Jian Liang, Xianda Guo +14
Multi-modal Large Language Models (MLLMs) integrate visual and linguistic reasoning to address complex tasks such as image captioning and visual question answering. While MLLMs dem…