activity
20242026
collaborators

17 papers

cs.CV2026

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting

Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5

Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…

cs.CV2026

LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension

Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3

Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…

cs.AI2026

Autoregressive Direct Preference Optimization

Masanari Oi, Mahiro Ukai, Masahiro Kaneko +2

Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However, the widespread reliance on the r…

cs.CV2026

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models

Masanari Oi, Koki Maeda, Ryuto Koike +3

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of inform…

cs.LG2026

Free Random Projection for In-Context Reinforcement Learning

Tomohiro Hayase, Benoît Collins, Nakamasa Inoue

Hierarchical inductive biases are hypothesized to promote generalizable policies in reinforcement learning, as demonstrated by explicit hyperbolic latent representations and archit…

cs.CV2026

BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment

Risa Shinoda, Kaede Shiohara, Nakamasa Inoue +3

Understanding animal species from multimodal data poses an emerging challenge at the intersection of computer vision and ecology. While recent biological models, such as BioCLIP, h…