109 citations · 256 across the 29 of their papers we have counts for
14 papers · 1 filter
How2: A Large-scale Dataset for Multimodal Language Understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar +4
In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequen…
Learning from Multiview Correlations in Open-Domain Videos
Nils Holzenberger, Shruti Palaskar, Pranava Madhyastha +2
An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views…
Multimodal Grounding for Sequence-to-Sequence Speech Recognition
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar +2
Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/…
Connectionist Temporal Localization for Sound Event Detection with Sequential Labeling
Yun Wang, Florian Metze
Research on sound event detection (SED) with weak labeling has mostly focused on presence/absence labeling, which provides no temporal information at all about the event occurrence…
A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling
Yun Wang, Juncheng Li, Florian Metze
Sound event detection (SED) entails two subtasks: recognizing what types of sound events are present in an audio stream (audio tagging), and pinpointing their onset and offset time…
Activity Recognition on a Large Scale in Short Videos - Moments in Time Dataset
Ankit Shah, Harini Kesavamoorthy, Poorva Rane +3
Moments capture a huge part of our lives. Accurate recognition of these moments is challenging due to the diverse and complex interpretation of the moments. Action recognition refe…