1 paper
Zeyuan Yang, Xueyang Yu, Delin Chen +2
Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand v…