activity
20242026
collaborators

6 papers

cs.CV2026

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

Yeeun Choi, Youngbeom Yoo, Joon-Young Lee +2

When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necess…

cs.AI2026

Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking

Jiwan Chung, JiHyuk Byun, Vibhav Vineet +1

Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improv…

cs.RO2026

Hierarchical Latent Action Model

Hanjung Kim, Lerrel Pinto, Seon Joo Kim

Latent Action Models (LAMs) enable learning from actionless data for applications ranging from robotic control to interactive world models. However, existing LAMs typically focus o…

cs.RO2025

UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations

Hanjung Kim, Jaehyun Kang, Hyolim Kang +3

Mimicry is a fundamental learning mechanism in humans, enabling individuals to learn new tasks by observing and imitating experts. However, applying this ability to robots presents…

cs.CV2025

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

Hyolim Kang, Yunsu Park, Youngbeom Yoo +2

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datas…

cs.CV2024

Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization

Jeongseok Hyun, Su Ho Han, Hyolim Kang +2

The vocabulary size in temporal action localization (TAL) is limited by the scarcity of large-scale annotated datasets. To overcome this, recent works integrate vision-language mod…