3 papers
cs.CV2025
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
Umihiro Kamoto, Tatsuya Ishibashi, Noriyuki Kugo
In this report, we present the winning solution that achieved the 1st place in the Complex Video Reasoning & Robustness Evaluation Challenge 2025. This challenge evaluates the abil…
cs.CV2025
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Noriyuki Kugo, Xiang Li, Zixin Li +9
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. Howe…
cs.CV2024
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
Noriyuki Kugo, Tatsuya Ishibashi, Kosuke Ono +1
This technical report provides a detailed description of our approach to the EgoSchema Challenge 2024. The EgoSchema Challenge aims to identify the most appropriate responses to qu…