1 paper
Haiyang Yan, Hongyun Zhou, Peng Xu +2
Despite rapid developments and widespread applications of MLLM agents, they still struggle with long-form video understanding (LVU) tasks, which are characterized by high informati…