1 paper
Xiangyu Zeng, Zhiqiu Zhang, Yuhan Zhu +12
Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse y…