1 paper
Chenhao Qiu, Yechao Zhang, Xin Luo +2
Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long vi…