3 papers
cs.CL2021
Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
Ramon Sanabria, Austin Waters, Jason Baldridge
Speech-based image retrieval has been studied as a proxy for joint representation learning, usually without emphasis on retrieval itself. As such, it is unclear how well speech-bas…
cs.CL2020
Crisscrossed Captions: Extended Intramodal and Intermodal Semantic Similarity Judgments for MS-COCO
Zarana Parekh, Jason Baldridge, Daniel Cer +2
By supporting multi-modal retrieval training and evaluation, image captioning datasets have spurred remarkable progress on representation learning. Unfortunately, datasets have lim…
eess.AS2018
From Audio to Semantics: Approaches to end-to-end spoken language understanding
Parisa Haghani, Arun Narayanan, Michiel Bacchiani +6
Conventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural languag…