17 papers
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting
Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5
Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…
Autoregressive Direct Preference Optimization
Masanari Oi, Mahiro Ukai, Masahiro Kaneko +2
Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However, the widespread reliance on the r…
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
Masanari Oi, Koki Maeda, Ryuto Koike +3
While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of inform…
Free Random Projection for In-Context Reinforcement Learning
Tomohiro Hayase, Benoît Collins, Nakamasa Inoue
Hierarchical inductive biases are hypothesized to promote generalizable policies in reinforcement learning, as demonstrated by explicit hyperbolic latent representations and archit…
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
Risa Shinoda, Kaede Shiohara, Nakamasa Inoue +3
Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, h…