1 paper
Hong Gao, Yiming Bao, Xuezhen Tu +3
Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume…