15 citations · 27 across the 11 of their papers we have counts for
7 papers · 1 filter
On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer
Liang Lu, Zhong Meng, Naoyuki Kanda +2
Hybrid Autoregressive Transducer (HAT) is a recently proposed end-to-end acoustic model that extends the standard Recurrent Neural Network Transducer (RNN-T) for the purpose of the…
Minimum Latency Training Strategies for Streaming Sequence-to-Sequence ASR
Hirofumi Inaguma, Yashesh Gaur, Liang Lu +2
Recently, a few novel streaming attention-based sequence-to-sequence (S2S) models have been proposed to perform online speech recognition with linear-time decoding complexity. Howe…
Exploring Pre-training with Alignments for RNN Transducer based End-to-End Speech Recognition
Hu Hu, Rui Zhao, Jinyu Li +2
Recently, the recurrent neural network transducer (RNN-T) architecture has become an emerging trend in end-to-end automatic speech recognition research due to its advantages of bei…
Semantic Mask for Transformer based End-to-End Speech Recognition
Chengyi Wang, Yu Wu, Yujiao Du +7
Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks. This approach takes advantage of t…
PyKaldi2: Yet another speech toolkit based on Kaldi and PyTorch
Liang Lu, Xiong Xiao, Zhuo Chen +1
We introduce PyKaldi2 speech recognition toolkit implemented based on Kaldi and PyTorch. While similar toolkits are available built on top of the two, a key feature of PyKaldi2 is…
Toward Computation and Memory Efficient Neural Network Acoustic Models with Binary Weights and Activations
Liang Lu
Neural network acoustic models have significantly advanced state of the art speech recognition over the past few years. However, they are usually computationally expensive due to t…