1 paper
Sungheon Jeong, Ryozo Masukawa, Jihong Park +5
While recent Large Vision-Language Models (LVLMs) exhibit strong multimodal reasoning abilities, they often produce ungrounded or hallucinated responses because they rely too heavi…