4 papers
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
Yeeun Choi, Youngbeom Yoo, Joon-Young Lee +2
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necess…
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
Hanjung Kim, Jaehyun Kang, Hyolim Kang +3
Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents…
Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
Hyolim Kang, Yunsu Park, Youngbeom Yoo +2
We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datas…
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
Jeongseok Hyun, Su Ho Han, Hyolim Kang +2
The vocabulary size in temporal action localization (TAL) is limited by the scarcity of large-scale annotated datasets. To overcome this, recent works integrate vision-language mod…