2 citations · 4 across the 16 of their papers we have counts for
1 paper · 1 filter
Yixin Ji, Fanghua Ye, Juntao Li +5
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing…