most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 184 across the 1 of their papers we have counts for

collaborators

5 papers

eess.AS2020

The RWTH ASR System for TED-LIUM Release 2: Improving Hybrid HMM with SpecAugment

Wei Zhou, Wilfried Michel, Kazuki Irie +3

We present a complete training pipeline to build a state-of-the-art hybrid HMM-based ASR system on the 2nd release of the TED-LIUM corpus. Data augmentation using SpecAugment is su…

cs.CL2019

Language Modeling with Deep Transformers

Kazuki Irie, Albert Zeyer, Ralf Schlüter +1

We explore deep autoregressive Transformer models in language modeling for speech recognition. We focus on two aspects. First, we revisit Transformer model configurations specifica…

cs.CL2019

RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation

Christoph Lüscher, Eugen Beck, Kazuki Irie +5

We present state-of-the-art automatic speech recognition (ASR) systems employing a standard hybrid DNN/HMM architecture compared to an attention-based encoder-decoder design for th…

cs.LG2019184 cited

Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

Jonathan Shen, Patrick Nguyen, Yonghui Wu +88

Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models a…

cs.CL2019

On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition

Kazuki Irie, Rohit Prabhavalkar, Anjuli Kannan +3

In conventional speech recognition, phoneme-based models outperform grapheme-based models for non-phonetic languages such as English. The performance gap between the two typically…