25 citations · 25 across the 1 of their papers we have counts for
6 papers · 1 filter
Multimodal Abstractive Summarization for How2 Videos
Shruti Palaskar, Jindrich Libovický, Spandana Gella +1
In this paper, we study abstractive summarization for open-domain videos. Unlike the traditional text news summarization, the goal is less to "compress" text information but rather…
Learned In Speech Recognition: Contextual Acoustic Word Embeddings
Shruti Palaskar, Vikas Raunak, Florian Metze
End-to-end acoustic-to-word speech recognition models have recently gained popularity because they are easy to train, scale well to large amounts of training data, and do not requi…
How2: A Large-scale Dataset for Multimodal Language Understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar +4
In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequen…
Multimodal Grounding for Sequence-to-Sequence Speech Recognition
Ozan Caglayan, Ramon Sanabria, Shruti Palaskar +2
Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/…
Linguistic unit discovery from multi-modal inputs in unwritten languages: Summary of the "Speaking Rosetta" JSALT 2017 Workshop
Odette Scharenborg, Laurent Besacier, Alan Black +16
We summarize the accomplishments of a multi-disciplinary workshop exploring the computational and scientific issues surrounding the discovery of linguistic units (subwords and word…
Combining LSTM and Latent Topic Modeling for Mortality Prediction
Yohan Jo, Lisa Lee, Shruti Palaskar
There is a great need for technologies that can predict the mortality of patients in intensive care units with both high accuracy and accountability. We present joint end-to-end ne…