2 papers
cs.CV2024
Streaming Dense Video Captioning
Xingyi Zhou, Anurag Arnab, Shyamal Buch +5
An ideal model for dense video captioning -- predicting captions localized temporally in a video -- should be able to handle long input videos, predict rich, detailed textual descr…
cs.CV2023
Open-Vocabulary Temporal Action Detection with Off-the-Shelf Image-Text Features
Vivek Rathod, Bryan Seybold, Sudheendra Vijayanarasimhan +4
Detecting actions in untrimmed videos should not be limited to a small, closed set of classes. We present a simple, yet effective strategy for open-vocabulary temporal action detec…