2 citations · 2 across the 6 of their papers we have counts for
4 papers · 1 filter
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM
Jiawei Yu, Yuang Li, Xiaosong Qiao +6
Text-to-speech (TTS) models have been widely adopted to enhance automatic speech recognition (ASR) systems using text-only corpora, thereby reducing the cost of labeling real speec…
Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
Yuang Li, Xiaosong Qiao, Xiaofeng Zhao +4
Large language models can enhance automatic speech recognition systems through generative error correction. In this paper, we propose Pinyin-enhanced GEC, which leverages Pinyi, th…
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
Jiaxin Guo, Minghan Wang, Xiaosong Qiao +9
Error correction techniques have been used to refine the output sentences from automatic speech recognition (ASR) models and achieve a lower word error rate (WER). Previous works u…
Using Large Language Model for End-to-End Chinese ASR and NER
Yuang Li, Jiawei Yu, Min Zhang +6
Mapping speech tokens to the same feature space as text tokens has become the paradigm for the integration of speech modality into decoder-only large language models (LLMs). An alt…