1 paper
Jiwoo Ha, Jongwoo Baek, Jinhyun So
Recent Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across various multimodal tasks that require understanding both visual and linguistic inputs. H…