collaborators

6 papers

cs.CV2026

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting

Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5

Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…

cs.CV2026

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…

cs.AI2026

Autoregressive Direct Preference Optimization

Masanari Oi, Mahiro Ukai, Masahiro Kaneko +2

Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However, the widespread reliance on the r…

cs.CV2026

DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning

Nakamasa Inoue, Kanoko Goto, Masanari Oi +4

Large vision-language models (LVLMs) have shown impressive performance across a broad range of multimodal tasks. However, robust image caption evaluation using LVLMs remains challe…

cs.CV2025

STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models

Mahiro Ukai, Shuhei Kurita, Nakamasa Inoue

Object state recognition aims to identify the specific condition of objects, such as their positional states (e.g., open or closed) and functional states (e.g., on or off). While r…

cs.CV2025

Referring Expression Comprehension for Small Objects

Kanoko Goto, Takumi Hirose, Mahiro Ukai +2

Referring expression comprehension (REC) aims to localize the target object described by a natural language expression. Recent advances in vision-language learning have led to sign…