13 papers
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting
Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5
Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
Koya Sakamoto, Taiki Miyanishi, Daichi Azuma +6
Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI…
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change…
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
Hisayuki Yokomizo, Taiki Miyanishi, Yan Gang +3
Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limi…
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
Daichi Yashima, Shuhei Kurita, Yusuke Oda +1
While multimodal large language models (MLLMs) have shown remarkable success across a wide range of tasks, long-form video understanding remains a significant challenge. In this st…