9 papers
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting
Fumihiko Tsuchiya, Taiki Miyanishi, Shunsuke Yasuki +5
Final-answer video QA can show whether a model predicts the right number, but not which instances it counted, when the supporting evidence occurs, or why it failed. We diagnose lon…
LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
Shunya Kato, Taiki Miyanishi, Shuhei Kurita +3
Egocentric videos capture rich and diverse human-object interactions and have emerged as a fundamental resource for understanding human activities related to objects. In this conte…
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
Hina Kogure, Kei Katsumata, Taiki Miyanishi +1
Dynamic urban environments are often captured by cameras placed at spatially separated locations with little or no view overlap. However, most existing 4D reconstruction methods as…
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
Koya Sakamoto, Taiki Miyanishi, Daichi Azuma +6
Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI…
NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change…
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
Hisayuki Yokomizo, Taiki Miyanishi, Yan Gang +3
Vision-Language Models (VLMs) are increasingly applied to robotic perception and manipulation, yet their ability to infer physical properties required for manipulation remains limi…