1 paper
Hongyu Qu, Guangming Yao, Ling Xing +7
Streaming video understanding requires multimodal large language models (MLLMs) to process continuous visual inputs and respond to user queries under strict causality and bounded m…