1 paper
Haopeng Jin, Hongzhu Yi, Wenlong Zhao +6
Long-video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame…