6 papers
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting
Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5
Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…
Autoregressive Direct Preference Optimization
Masanari Oi, Mahiro Ukai, Masahiro Kaneko +2
Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However, the widespread reliance on the r…
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
Nakamasa Inoue, Kanoko Goto, Masanari Oi +4
Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challe…
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue
Object state recognition aims to identify the specific condition of objects, such as their positional states (e.g., open or closed) and functional states (e.g., on or off). While r…
Referring Expression Comprehension for Small Objects
Kanoko Goto, Takumi Hirose, Mahiro Ukai +2
Referring expression comprehension (REC) aims to localize the target object described by a natural language expression. Recent advances in vision-language learning have led to sign…