7 citations · 40 across the 58 of their papers we have counts for
1 paper · 1 filter
Kejian Zhu, Zhuoran Jin, Hongbang Yuan +6
The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. How…