Showing 2026 · cs.CVShow all
2 papers · 2 filters
cs.CV2026
Rigel: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
Shuitsu Koyama, Kazuki Matsuda, Yuiga Wada +3
Automatic evaluation of image and video captioning is essential for benchmarking multimodal systems, although standard evaluation metrics show limited alignment with human judgment…
cs.CV2026
MLLM-as-a-Judge Exhibits Model Preference Bias
Shuitsu Koyama, Yuiga Wada, Daichi Yashima +1
Automatic evaluation using multimodal large language models (MLLMs), commonly referred to as MLLM-as-a-Judge, has been widely used to measure model performance. If such MLLM-as-a-J…