25 citations · 25 across the 1 of their papers we have counts for
6 papers · 1 filter
How2: A Large-scale Dataset for Multimodal Language Understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar +4
In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequen…
Learning from Multiview Correlations in Open-Domain Videos
Nils Holzenberger, Shruti Palaskar, Pranava Madhyastha +2
An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views…
Multimodal Grounding for Sequence-to-Sequence Speech Recognition
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar +2
Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/…
Acoustic-to-Word Recognition with Sequence-to-Sequence Models
Shruti Palaskar, Florian Metze
Acoustic-to-Word recognition provides a straightforward solution to end-to-end speech recognition without needing external decoding, language model re-scoring or lexicon. While cha…
End-to-End Multimodal Speech Recognition
Shruti Palaskar, Ramon Sanabria, Florian Metze
Transcription or sub-titling of open-domain videos is still a challenging domain for Automatic Speech Recognition (ASR) due to the data's challenging acoustics, variable signal pro…
Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop
Odette Scharenborg, Laurent Besacier, Alan Black +16
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and word…