5 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CV2024
TWLV-I: Analysis and Insights from Holistic Evaluation on Video Foundation Models
Hyeongmin Lee, Jin-Young Kim, Kyungjune Baek +18
In this work, we discuss evaluating video foundation models in a fair and robust manner. Unlike language or image foundation models, many video foundation models are evaluated with…
cs.MM2024
Pegasus-v1 Technical Report
Raehyuk Jung, Hyojun Go, Jaehyuk Yi +41
This technical report introduces Pegasus-1, a multimodal language model specialized in video content understanding and interaction through natural language. Pegasus-1 is designed t…
cs.CV2023★ 5 cited
Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment
Yongrae Jo, Seongyun Lee, Aiden SJ Lee +3
Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments pa…