1 paper
Seojeong Park, Jiho Choi, Junyong Kang +3
Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visua…