1 paper
Jaeha Choi, Jin Won Lee, Siwoo You +1
Advances in vision-language models (VLMs) have achieved remarkable success on complex multimodal reasoning tasks, leading to the assumption that they should also excel at reading a…