4 papers
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Yeeun Choi, Youngbeom Yoo, Joon-Young Lee +2
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necess…
FUSE: Ensembling Verifiers with Zero Labeled Data
Joonhyuk Lee, Virginia Ma, Sarah Zhao +4
Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves u…
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
Jeongseok Hyun, Sukjun Hwang, Su Ho Han +6
Video large language models (LLMs) achieve strong video understanding by leveraging a large number of spatio-temporal tokens, but suffer from quadratic computational scaling with t…
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
Jeongseok Hyun, Su Ho Han, Hyolim Kang +2
The vocabulary size in temporal action localization (TAL) is limited by the scarcity of large-scale annotated datasets. To overcome this, recent works integrate vision-language mod…