1 paper · 1 filter
Wenjin Liu, Haoran Luo, Fayuan Ke +5
Recently, multimodal large language models (MLLMs) have demonstrated strong potential in visual understanding and complex reasoning tasks. However, existing methods often struggle…