1 paper · 1 filter
Jeong Hun Yeo, Sangyun Chung, Sungjune Park +3
Long-video understanding remains a significant challenge for Multimodal Large Language Models (MLLMs) due to inherent token limitations and the complexity of capturing long-term te…