4 papers
LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
Shinnosuke Hirano, Yuiga Wada, Kazuki Matsuda +2
We focus on the automatic evaluation of image captions in both reference-based and reference-free settings. Existing metrics based on large language models (LLMs) favor their own g…
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
Kazuki Matsuda, Yuiga Wada, Shinnosuke Hirano +2
In this study, we focus on the automatic evaluation of long and detailed image captions generated by multimodal Large Language Models (MLLMs). Most existing automatic evaluation me…
Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
Shinnosuke Hirano, Yuiga Wada, Tsumugi Iida +1
In this study, we consider the problem of generating visual explanations in visual foundation models. Numerous methods have been proposed for this purpose; however, they often cann…
Capturing Fine-Grained Alignments Improves 3D Affordance Detection
Junsei Tokumitsu, Yuiga Wada
In this work, we address the challenge of affordance detection in 3D point clouds, a task that requires effectively capturing fine-grained alignments between point clouds and text.…