activity
20182024
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 195 across the 7 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2024

Deferred NAM: Low-latency Top-K Context Injection via Deferred Context Encoding for Non-Streaming ASR

Zelin Wu, Gan Song, Christopher Li +9

Contextual biasing enables speech recognizers to transcribe important phrases in the speaker's context, such as contact names, even if they are rare in, or absent from, the trainin…

cs.CL20231 cited

SLM: Bridge the thin gap between speech and text foundation models

Mingqiu Wang, Wei Han, Izhak Shafran +15

We present a joint Speech and Language Model (SLM), a multitask, multilingual, and dual-modal model that takes advantage of pretrained foundational speech and language models. SLM…

cs.CL2023

Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm

Weiran Wang, Zelin Wu, Diamantino Caseiro +10

Contextual biasing refers to the problem of biasing the automatic speech recognition (ASR) systems towards rare entities that are relevant to the specific user or application scena…

cs.CL20203 cited

A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency

Tara N. Sainath, Yanzhang He, Bo Li +26

Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e…

cs.CL20197 cited

Phoneme-Based Contextualization for Cross-Lingual Speech Recognition in End-to-End Models

Ke Hu, Antoine Bruguier, Tara N. Sainath +2

Contextual automatic speech recognition, i.e., biasing recognition towards a given context (e.g. user's playlists, or contacts), is challenging in end-to-end (E2E) models. Such mod…

cs.CL2018

Streaming End-to-end Speech Recognition For Mobile Devices

Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar +17

End-to-end (E2E) models, which directly predict output character sequences given input speech, are good candidates for on-device speech recognition. E2E models, however, present nu…