1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Kejian Zhu, Zhuoran Jin, Hongbang Yuan +6
The sequential structure of videos poses a challenge to the ability of multimodal large language models (MLLMs) to locate multi-frame evidence and conduct multimodal reasoning. How…