activity
20192022
most citedSelf-Attention Transducers for End-to-End Speech Recognition

85 citations · 116 across the 9 of their papers we have counts for

collaborators
Showing eess.ASShow all

9 papers · 1 filter

eess.AS2022

Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition

Ye Bai, Jie Li, Wenjing Han +5

While transformers and their variant conformers show promising performance in speech recognition, the parameterized property leads to much memory cost during training and inference…

eess.AS2021

FSR: Accelerating the Inference Process of Transducer-Based Models by Applying Fast-Skip Regularization

Zhengkun Tian, Jiangyan Yi, Ye Bai +3

Transducer-based models, such as RNN-Transducer and transformer-transducer, have achieved great success in speech recognition. A typical transducer model decodes the output sequenc…

eess.AS2020

One In A Hundred: Select The Best Predicted Sequence from Numerous Candidates for Streaming Speech Recognition

Zhengkun Tian, Jiangyan Yi, Ye Bai +3

The RNN-Transducers and improved attention-based encoder-decoder models are widely applied to streaming speech recognition. Compared with these two end-to-end models, the CTC model…

eess.AS202010 cited

Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition

Zhengkun Tian, Jiangyan Yi, Jianhua Tao +3

Non-autoregressive transformer models have achieved extremely fast inference speed and comparable performance with autoregressive sequence-to-sequence models in neural machine tran…

eess.AS20205 cited

Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition

Ye Bai, Jiangyan Yi, Jianhua Tao +3

Although attention based end-to-end models have achieved promising performance in speech recognition, the multi-pass forward computation in beam-search increases inference time cos…

eess.AS20195 cited

Synchronous Transformers for End-to-End Speech Recognition

Zhengkun Tian, Jiangyan Yi, Ye Bai +3

For most of the attention-based sequence-to-sequence models, the decoder predicts the output sequence conditioned on the entire input sequence processed by the encoder. The asynchr…