99 citations · 123 across the 11 of their papers we have counts for
20 papers · 1 filter
Learning Human Action Recognition Representations Without Real Humans
Howard Zhong, Samarth Mishra, Donghyun Kim +7
Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets…
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
Nina Shvetsova, Anna Kukleva, Xudong Hong +3
Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR…
In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval
Nina Shvetsova, Anna Kukleva, Bernt Schiele +1
Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, when transferring them to the task of video retrieva…
Preserving Modality Structure Improves Multi-Modal Learning
Swetha Sirnam, Mamshad Nayeem Rizve, Nina Shvetsova +2
Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human…
Learning Situation Hyper-Graphs for Video Question Answering
Aisha Urooj Khan, Hilde Kuehne, Bo Wu +5
Answering questions about complex situations in videos requires not only capturing the presence of actors, objects, and their relations but also the evolution of these relationship…
WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
Marius Bock, Hilde Kuehne, Kristof Van Laerhoven +1
Research has shown the complementarity of camera- and inertial-based data for modeling human activities, yet datasets with both egocentric video and inertial-based sensor data rema…