4 papers
HydraFormer: One Encoder For All Subsampling Rates
Yaoxun Xu, Xingchen Song, Zhiyong Wu +3
In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situati…
Spike-Triggered Contextual Biasing for End-to-End Mandarin Speech Recognition
Kaixun Huang, Ao Zhang, Binbin Zhang +3
The attention-based deep contextual biasing method has been demonstrated to effectively improve the recognition performance of end-to-end automatic speech recognition (ASR) systems…
LightGrad: Lightweight Diffusion Probabilistic Model for Text-to-Speech
Jie Chen, Xingchen Song, Zhendong Peng +3
Recent advances in neural text-to-speech (TTS) models bring thousands of TTS applications into daily life, where models are deployed in cloud to provide services for customs. Among…
CB-Conformer: Contextual biasing Conformer for biased word recognition
Yaoxun Xu, Baiji Liu, Qiaochu Huang and +4
Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in…