activity
20162022
most citedMasked Autoencoders that Listen

109 citations · 256 across the 29 of their papers we have counts for

collaborators
Showing 2018Show all

14 papers · 1 filter

cs.CL2018

How2: A Large-scale Dataset for Multimodal Language Understanding

Ramon Sanabria, Ozan Caglayan, Shruti Palaskar +4

In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequen…

cs.LG2018

Learning from Multiview Correlations in Open-Domain Videos

Nils Holzenberger, Shruti Palaskar, Pranava Madhyastha +2

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views…

cs.CL2018

Multimodal Grounding for Sequence-to-Sequence Speech Recognition

Ozan Caglayan, Ramon Sanabria, Shruti Palaskar +2

Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/…

cs.SD2018

Connectionist Temporal Localization for Sound Event Detection with Sequential Labeling

Yun Wang, Florian Metze

Research on sound event detection (SED) with weak labeling has mostly focused on presence/absence labeling, which provides no temporal information at all about the event occurrence…

cs.SD2018

A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling

Yun Wang, Juncheng Li, Florian Metze

Sound event detection (SED) entails two subtasks: recognizing what types of sound events are present in an audio stream (audio tagging), and pinpointing their onset and offset time…

cs.CV2018

Activity Recognition on a Large Scale in Short Videos - Moments in Time Dataset

Ankit Shah, Harini Kesavamoorthy, Poorva Rane +3

Moments capture a huge part of our lives. Accurate recognition of these moments is challenging due to the diverse and complex interpretation of the moments. Action recognition refe…