1 paper · 1 filter
Haozhe Qi, Kevin Qu, Mahdi Rad +3
Long video understanding remains challenging for Multi-modal Large Language Models (MLLMs) due to high memory costs and context-length limits. Prior approaches mitigate this by sco…