1 paper
Jin Cui, Xinyue Long, Xunyong Zhang +5
Multimodal Large Language Models (MLLMs) have made remarkable progress on vision-language reasoning, yet most methods still compress visual evidence into discrete textual thoughts,…