67 citations · 89 across the 15 of their papers we have counts for
5 papers · 1 filter
Speech Corpora Divergence Based Unsupervised Data Selection for ASR
Changfeng Gao, Gaofeng Cheng, Pengyuan Zhang +1
Selecting application scenarios matching data is important for the automatic speech recognition (ASR) training, but it is difficult to measure the matching degree of the training c…
Summary on the ISCSLP 2022 Chinese-English Code-Switching ASR Challenge
Shuhao Deng, Chengfei Li, Jinfeng Bai +6
Code-switching automatic speech recognition becomes one of the most challenging and the most valuable scenarios of automatic speech recognition, due to the code-switching phenomeno…
Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset
Zehui Yang, Yifan Chen, Lei Luo +9
This paper introduces a high-quality rich annotated Mandarin conversational (RAMC) speech dataset called MagicData-RAMC. The MagicData-RAMC corpus contains 180 hours of conversatio…
Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable
Ruiliu Fu, Han Wang, Xuejun Zhang +2
Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the machine reasoning pr…
Reminding the Incremental Language Model via Data-Free Self-Distillation
Han Wang, Ruiliu Fu, Chengzhang Li +3
Incremental language learning with pseudo-data can alleviate catastrophic forgetting in neural networks. However, to obtain better performance, former methods have higher demands f…