activity
20162023
most citedSelf-Attention Transducers for End-to-End Speech Recognition

85 citations · 472 across the 51 of their papers we have counts for

collaborators
Showing 2019 · eess.ASShow all

7 papers · 2 filters

eess.AS2019★ 5 cited

Synchronous Transformers for End-to-End Speech Recognition

Zhengkun Tian, Jiangyan Yi, Ye Bai +3

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchr…

eess.AS2019

Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data

Ye Bai, Jiangyan Yi, Jianhua Tao +3

Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained…

eess.AS2019★ 7 cited

Domain adversarial learning for emotion recognition

Zheng Lian, Jianhua Tao, Bin Liu +1

In practical applications for emotion recognition, users do not always exist in the training corpus. The mismatch between training speakers and testing speakers affects the perform…

eess.AS2019★ 17 cited

Unsupervised Representation Learning with Future Observation Prediction for Speech Emotion Recognition

Zheng Lian, Jianhua Tao, Bin Liu +1

Prior works on speech emotion recognition utilize various unsupervised learning approaches to deal with low-resource samples. However, these methods pay less attention to modeling…

eess.AS2019★ 85 cited

Self-Attention Transducers for End-to-End Speech Recognition

Zhengkun Tian, Jiangyan Yi, Jianhua Tao +2

Recurrent neural network transducers (RNN-T) have been successfully applied in end-to-end speech recognition. However, the recurrent structure makes it difficult for parallelizatio…

eess.AS2019★ 2 cited

Forward-Backward Decoding for Regularizing End-to-End TTS

Yibin Zheng, Xi Wang, Lei He +4

Neural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scal…