1 paper · 1 filter
Saul Santos, António Farinhas, Daniel C. McNamee +1
Current video-language models struggle with long-video understanding due to limited context lengths and reliance on sparse frame subsampling, often leading to information loss. Thi…