1 paper · 1 filter
Jiwan Chung, Junhyeok Kim, Siyeol Kim +3
When thinking with images, humans rarely rely on a single glance: they revisit visual evidence while reasoning. In contrast, most Multimodal Language Models encode an image once to…