activity
20182020
most citedStreaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition

13 citations · 41 across the 4 of their papers we have counts for

collaborators

9 papers

cs.SD202010 cited

Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model

Zhifu Gao, Shiliang Zhang, Ming Lei +1

Recently, online end-to-end ASR has gained increasing attention. However, the performance of online systems still lags far behind that of offline systems, with a large gap in quali…

eess.AS2020

DeviceTTS: A Small-Footprint, Fast, Stable Network for On-Device Text-to-Speech

Zhiying Huang, Hao Li, Ming Lei

With the number of smart devices increasing, the demand for on-device text-to-speech (TTS) increases rapidly. In recent years, many prominent End-to-End TTS methods have been propo…

cs.SD20209 cited

SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition

Zhifu Gao, Shiliang Zhang, Ming Lei +1

End-to-end speech recognition has become popular in recent years, since it can integrate the acoustic, pronunciation and language models into a single neural network. Among end-to-…

cs.SD202013 cited

Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition

Shiliang Zhang, Zhifu Gao, Haoneng Luo +4

Recently, streaming end-to-end automatic speech recognition (E2E-ASR) has gained more and more attention. Many efforts have been paid to turn the non-streaming attention-based E2E-…

cs.SD2020

Simplified Self-Attention for Transformer-based End-to-End Speech Recognition

Haoneng Luo, Shiliang Zhang, Ming Lei +1

Transformer models have been introduced into end-to-end speech recognition with state-of-the-art performance on various tasks owing to their superiority in modeling long-term depen…

eess.AS20199 cited

Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition

Shiliang Zhang, Ming Lei, Zhijie Yan

Connectionist Temporal Classification (CTC) based end-to-end speech recognition system usually need to incorporate an external language model by using WFST-based decoding in order…