1 paper
Ye Wang, Qianglong Chen, Zejun Li +4
Multimodal Large Language Models (MLLMs) have shown impressive performance on vision-language tasks, but their long Chain-of-Thought (CoT) capabilities in multimodal scenarios rema…