1 paper
Mengqi Shi, Haopeng Zhang
While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains…