1 paper
Dong-Hee Kim, Seonwoo Choi, Changbeen Kim +6
Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, whil…