activity
20162022
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 195 across the 6 of their papers we have counts for

collaborators

8 papers

cs.CL2022

Neural-FST Class Language Model for End-to-End Speech Recognition

Antoine Bruguier, Duc Le, Rohit Prabhavalkar +7

We propose Neural-FST Class Language Model (NFCLM) for end-to-end speech recognition, a novel method that combines neural network language models (NNLMs) and finite state transduce…

cs.CL20201 cited

Improved Neural Language Model Fusion for Streaming Recurrent Neural Network Transducer

Suyoun Kim, Yuan Shangguan, Jay Mahadeokar +4

Recurrent Neural Network Transducer (RNN-T), like most end-to-end speech recognition model architectures, has an implicit neural network language model (NNLM) and cannot easily lev…

cs.CL20203 cited

A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency

Tara N. Sainath, Yanzhang He, Bo Li +26

Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e…

cs.CL20197 cited

Phoneme-Based Contextualization for Cross-Lingual Speech Recognition in End-to-End Models

Ke Hu, Antoine Bruguier, Tara N. Sainath +2

Contextual automatic speech recognition, i.e., biasing recognition towards a given context (e.g. user's playlists, or contacts), is challenging in end-to-end (E2E) models. Such mod…

cs.LG2019184 cited

Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

Jonathan Shen, Patrick Nguyen, Yonghui Wu +88

Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models a…

cs.CL2019

On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition

Kazuki Irie, Rohit Prabhavalkar, Anjuli Kannan +3

In conventional speech recognition, phoneme-based models outperform grapheme-based models for non-phonetic languages such as English. The performance gap between the two typically…