1 paper
Minkuk Kim, Suyong Yun, Young Tae Kim +3
Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where u…