4 papers
Revisiting Interpolation Augmentation for Speech-to-Text Generation
Chen Xu, Jie Wang, Xiaoqian Liu +6
Speech-to-text (S2T) generation systems frequently face challenges in low-resource scenarios, primarily due to the lack of extensive labeled datasets. One emerging solution is cons…
Bridging the Gaps of Both Modality and Language: Synchronous Bilingual CTC for Speech Translation and Speech Recognition
Chen Xu, Xiaoqian Liu, Erfeng He +6
In this study, we present synchronous bilingual Connectionist Temporal Classification (CTC), an innovative framework that leverages dual CTC to bridge the gaps of both modality and…
CTC-based Non-autoregressive Speech Translation
Chen Xu, Xiaoqian Liu, Xiaowen Liu +9
Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation a…
Bridging the Granularity Gap for Acoustic Modeling
Chen Xu, Yuhao Zhang, Chengbo Jiao +7
While Transformer has become the de-facto standard for speech, modeling upon the fine-grained frame-level features remains an open challenge of capturing long-distance dependencies…