1 paper
Kaishen Wang, Dongdi Zhao, Yijun Liang +4
Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the…