activity
20162022
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 230 across the 16 of their papers we have counts for

collaborators

29 papers

cs.CL20221 cited

JOIST: A Joint Speech and Text Streaming Model For ASR

Tara N. Sainath, Rohit Prabhavalkar, Ankur Bapna +6

We present JOIST, an algorithm to train a streaming, cascaded, encoder end-to-end (E2E) model with both speech-text paired inputs, and text-only unpaired inputs. Unlike previous wo…

cs.CL2022

Neural-FST Class Language Model for End-to-End Speech Recognition

Antoine Bruguier, Duc Le, Rohit Prabhavalkar +7

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transduce…

eess.AS2021

A Neural Acoustic Echo Canceller Optimized Using An Automatic Speech Recognizer And Large Scale Synthetic Data

Nathan Howard, Alex Park, Turaj Zakizadeh Shabestary +2

We consider the problem of recognizing speech utterances spoken to a device which is generating a known sound waveform; for example, recognizing queries issued to a digital assista…

cs.CL2021

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

Yangyang Shi, Varun Nagaraja, Chunyang Wu +9

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…

cs.SD2021

Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

Yuan Shangguan, Rohit Prabhavalkar, Hang Su +8

As speech-enabled devices such as smartphones and smart speakers become increasingly ubiquitous, there is growing interest in building automatic speech recognition (ASR) systems th…

eess.AS20212 cited

Learning Word-Level Confidence For Subword End-to-End ASR

David Qiu, Qiujia Li, Yanzhang He +9

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed trainin…