activity
20172024
most citedA Hybrid RNN-HMM Approach for Weakly Supervised Temporal Action Segmentation

99 citations · 123 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

20 papers · 1 filter

cs.CV20231 cited

Learning Human Action Recognition Representations Without Real Humans

Howard Zhong, Samarth Mishra, Donghyun Kim +7

Pre-training on massive video datasets has become essential to achieve high action recognition performance on smaller downstream datasets. However, most large-scale video datasets…

cs.CV2023

HowToCaption: Prompting LLMs to Transform Video Annotations at Scale

Nina Shvetsova, Anna Kukleva, Xudong Hong +3

Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR…

cs.CV2023

In-Style: Bridging Text and Uncurated Videos with Style Transfer for Text-Video Retrieval

Nina Shvetsova, Anna Kukleva, Bernt Schiele +1

Large-scale noisy web image-text datasets have been proven to be efficient for learning robust vision-language models. However, when transferring them to the task of video retrieva…

cs.CV2023

Preserving Modality Structure Improves Multi-Modal Learning

Swetha Sirnam, Mamshad Nayeem Rizve, Nina Shvetsova +2

Self-supervised learning on large-scale multi-modal datasets allows learning semantically meaningful embeddings in a joint multi-modal representation space without relying on human…

cs.CV20234 cited

Learning Situation Hyper-Graphs for Video Question Answering

Aisha Urooj Khan, Hilde Kuehne, Bo Wu +5

Answering questions about complex situations in videos requires not only capturing the presence of actors, objects, and their relations but also the evolution of these relationship…

cs.CV2023

WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition

Marius Bock, Hilde Kuehne, Kristof Van Laerhoven +1

Research has shown the complementarity of camera- and inertial-based data for modeling human activities, yet datasets with both egocentric video and inertial-based sensor data rema…