1 paper · 1 filter
Kunhao Li, Wenhao Li, Di Wu +4
Multimodal Large Language Models (MLLMs) extend foundation models to real-world applications by integrating inputs such as text and vision. However, their broad knowledge capacity…