most citedSLM: Bridge the thin gap between speech and text foundation models

1 citations · 2 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL20231 cited

SLM: Bridge the thin gap between speech and text foundation models

Mingqiu Wang, Wei Han, Izhak Shafran +15

We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM…

cs.CL2023

Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm

Weiran Wang, Zelin Wu, Diamantino Caseiro +10

Contextual biasing refers to the problem of biasing the automatic speech recognition (ASR) systems towards rare entities that are relevant to the specific user or application scena…

eess.AS20231 cited

Massive End-to-end Models for Short Search Queries

Weiran Wang, Rohit Prabhavalkar, Dongseong Hwang +11

In this work, we investigate two popular end-to-end automatic speech recognition (ASR) models, namely Connectionist Temporal Classification (CTC) and RNN-Transducer (RNN-T), for of…

cs.CL2023

Text Injection for Capitalization and Turn-Taking Prediction in Speech Models

Shaan Bijwadia, Shuo-yiin Chang, Weiran Wang +3

Text injection for automatic speech recognition (ASR), wherein unpaired text-only data is used to supplement paired audio-text data, has shown promising improvements for word error…

cs.CL2023

Improving Joint Speech-Text Representations Without Alignment

Cal Peyser, Zhong Meng, Ke Hu +5

The last year has seen astonishing progress in text-prompted image generation premised on the idea of a cross-modal representation space in which the text and image domains are rep…

eess.AS2023

JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition

Zhong Meng, Weiran Wang, Rohit Prabhavalkar +7

We propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves ra…