5 citations · 5 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
TWLV-I: Analysis and Insights from Holistic Evaluation on Video Foundation Models
Hyeongmin Lee, Jin-Young Kim, Kyungjune Baek +18
In this work, we discuss evaluating video foundation models in a fair and robust manner. Unlike language or image foundation models, many video foundation models are evaluated with…
cs.CV2023★ 5 cited
Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment
Yongrae Jo, Seongyun Lee, Aiden SJ Lee +3
Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments pa…