13 citations · 41 across the 4 of their papers we have counts for
9 papers
Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
Zhifu Gao, Shiliang Zhang, Ming Lei +1
Recently, online end-to-end ASR has gained increasing attention. However, the performance of online systems still lags far behind that of offline systems, with a large gap in quali…
DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech
Zhiying Huang, Hao Li, Ming Lei
With the number of smart devices increasing, the demand for on-device text-to-speech (TTS) increases rapidly. In recent years, many prominent End-to-End TTS methods have been propo…
SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition
Zhifu Gao, Shiliang Zhang, Ming Lei +1
End-to-end speech recognition has become popular in recent years, since it can integrate the acoustic, pronunciation and language models into a single neural network. Among end-to-…
Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition
Shiliang Zhang, Zhifu Gao, Haoneng Luo +4
Recently, streaming end-to-end automatic speech recognition (E2E-ASR) has gained more and more attention. Many efforts have been paid to turn the non-streaming attention-based E2E-…
Simplified Self-Attention for Transformer-based End-to-End Speech Recognition
Haoneng Luo, Shiliang Zhang, Ming Lei +1
Transformer models have been introduced into end-to-end speech recognition with state-of-the-art performance on various tasks owing to their superiority in modeling long-term depen…
Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
Shiliang Zhang, Ming Lei, Zhijie Yan
Connectionist Temporal Classification (CTC) based end-to-end speech recognition system usually need to incorporate an external language model by using WFST-based decoding in order…