1 paper · 1 filter
Ziyi Wang, Haoran Wu, Yiming Rong +5
Long video understanding is a complex task that requires both spatial detail and temporal awareness. While Vision-Language Models (VLMs) obtain frame-level understanding capabiliti…