1 paper · 1 filter
Junbo Zou, Ziheng Huang, Shengjie Zhang +2
Long-form video understanding remains challenging for Vision-Language Models (VLMs) due to the inherent tension between computational constraints and the need to capture informatio…