activity
20172022
most citedTransformer-Transducer: End-to-End Speech Recognition with Self-Attention

66 citations · 156 across the 21 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS20211 cited

Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution

Yangyang Shi, Chunyang Wu, Dilin Wang +9

This paper improves the streaming transformer transducer for speech recognition by using non-causal convolution. Many works apply the causal convolution to improve streaming transf…

eess.AS20211 cited

On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models

Xiaohui Zhang, Vimal Manohar, David Zhang +7

Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…

eess.AS20205 cited

Weak-Attention Suppression For Transformer Based Speech Recognition

Yangyang Shi, Yongqiang Wang, Chunyang Wu +5

Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acousti…

eess.AS201966 cited

Transformer-Transducer: End-to-End Speech Recognition with Self-Attention

Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar +6

We explore options to use Transformer networks in neural transducer for end-to-end speech recognition. Transformer networks use self-attention for sequence modeling and comes with…

eess.AS20199 cited

From Senones to Chenones: Tied Context-Dependent Graphemes for Hybrid Speech Recognition

Duc Le, Xiaohui Zhang, Weiyi Zheng +3

There is an implicit assumption that traditional hybrid approaches for automatic speech recognition (ASR) cannot directly model graphemes and need to rely on phonetic lexicons to g…

eess.AS20191 cited

G2G: TTS-Driven Pronunciation Learning for Graphemic Hybrid ASR

Duc Le, Thilo Koehler, Christian Fuegen +1

Grapheme-based acoustic modeling has recently been shown to outperform phoneme-based approaches in both hybrid and end-to-end automatic speech recognition (ASR), even on non-phonem…