3 papers
eess.AS2024
HydraFormer: One Encoder For All Subsampling Rates
Yaoxun Xu, Xingchen Song, Zhiyong Wu +3
In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situati…
cs.SD2023
Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation
Jiaxu Zhu, Weinan Tong, Yaoxun Xu +6
Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) pe…
cs.SD2023
CB-Conformer: Contextual biasing Conformer for biased word recognition
Yaoxun Xu, Baiji Liu, Qiaochu Huang and +4
Due to the mismatch between the source and target domains, how to better utilize the biased word information to improve the performance of the automatic speech recognition model in…