Showing cs.CLShow all
3 papers · 1 filter
cs.CL2024
High-precision Voice Search Query Correction via Retrievable Speech-text Embedings
Christopher Li, Gary Wang, Kyle Kastner +9
Automatic speech recognition (ASR) systems can suffer from poor recall for various reasons, such as noisy audio, lack of sufficient training data, etc. Previous work has shown that…
cs.CL2023
Improving Joint Speech-Text Representations Without Alignment
Cal Peyser, Zhong Meng, Ke Hu +5
The last year has seen astonishing progress in text-prompted image generation premised on the idea of a cross-modal representation space in which the text and image domains are rep…
cs.CL2023
Understanding Shared Speech-Text Representations
Gary Wang, Kyle Kastner, Ankur Bapna +4
Recently, a number of approaches to train speech models by incorpo-rating text into end-to-end models have been developed, with Mae-stro advancing state-of-the-art automatic speech…