1 paper
Hongcheng Gao, Zihao Huang, Jingyi Tang +12
A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language reasoning (e.g., chain-of-thought, in-co…