Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
Masanari Oi, Koki Maeda, Ryuto Koike +3
While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of inform…
cs.CV2026
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
Nakamasa Inoue, Kanoko Goto, Masanari Oi +4
Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challe…