1 paper
Woojun Jung, Jaehoon Go, Mingyu Jeon +2
Multimodal Large Language Models (MLLMs) demonstrate impressive reasoning capabilities, but often fail to perceive fine-grained visual details, limiting their applicability in prec…