5 papers
ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA
Minkuk Kim, Suyong Yun, Young Tae Kim +3
Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where u…
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
Geo Ahn, Jiwook Han, Youngrae Kim +2
Fine-tuning MLLMs for Video Temporal Grounding (VTG) often improves in-domain performance but degrades sharply under domain shift. In this work, we find that this failure is primar…
Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs
Jongseo Lee, Hyuntak Lee, Sunghun Kim +3
Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion directio…
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
Joohyun Chang, Soyeon Hong, Hyogun Lee +4
In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoin…
HiCM: Hierarchical Compact Memory Modeling for Dense Video Captioning
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon +2
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and local…