85 citations · 472 across the 51 of their papers we have counts for
7 papers · 2 filters
Synchronous Transformers for End-to-End Speech Recognition
Zhengkun Tian, Jiangyan Yi, Ye Bai +3
For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchr…
Integrating Knowledge into End-to-End Speech Recognition from External Text-Only Data
Ye Bai, Jiangyan Yi, Jianhua Tao +3
Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained…
Domain adversarial learning for emotion recognition
Zheng Lian, Jianhua Tao, Bin Liu +1
In practical applications for emotion recognition, users do not always exist in the training corpus. The mismatch between training speakers and testing speakers affects the perform…
Unsupervised Representation Learning with Future Observation Prediction for Speech Emotion Recognition
Zheng Lian, Jianhua Tao, Bin Liu +1
Prior works on speech emotion recognition utilize various unsupervised learning approaches to deal with low-resource samples. However, these methods pay less attention to modeling…
Self-Attention Transducers for End-to-End Speech Recognition
Zhengkun Tian, Jiangyan Yi, Jianhua Tao +2
Recurrent neural network transducers (RNN-T) have been successfully applied in end-to-end speech recognition. However, the recurrent structure makes it difficult for parallelizatio…
Forward-Backward Decoding for Regularizing End-to-End TTS
Yibin Zheng, Xi Wang, Lei He +4
Neural end-to-end TTS can generate very high-quality synthesized speech, and even close to human recording within similar domain text. However, it performs unsatisfactory when scal…