1 paper
Minghao Qin, Yan Shu, Peitian Zhang +6
Long-video understanding (LVU) remains a severe challenge for existing multimodal large language models (MLLMs), primarily due to the prohibitive computational cost. Recent approac…