2 papers
cs.CV2026
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
Minkuk Kim, Suyong Yun, Young Tae Kim +3
Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where u…
cs.CV2024
HiCM: Hierarchical Compact Memory Modeling for Dense Video Captioning
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon +2
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and local…