14 citations · 40 across the 14 of their papers we have counts for
25 papers
Cross-Modal Knowledge Distillation for Speech Large Language Models
Enzhi Wang, Qicheng Li, Zhiyuan Tang +1
In this work, we present the first systematic evaluation of catastrophic forgetting and modality inequivalence in speech large language models, showing that introducing speech capa…
Chain of Correction for Full-text Speech Recognition with Large Language Models
Zhiyuan Tang, Dong Wang, Zhikai Zhou +3
Full-text error correction with Large Language Models (LLMs) for Automatic Speech Recognition (ASR) is attracting increased attention for its ability to address a wide range of err…
Full-text Error Correction for Chinese Speech Recognition with Large Language Model
Zhiyuan Tang, Dong Wang, Shen Huang +1
Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from s…
Semantic Data Augmentation for End-to-End Mandarin Speech Recognition
Jianwei Sun, Zhiyuan Tang, Hengxin Yin +6
End-to-end models have gradually become the preferred option for automatic speech recognition (ASR) applications. During the training of end-to-end ASR, data augmentation is a quit…
Can We Trust Deep Speech Prior?
Ying Shi, Haolin Chen, Zhiyuan Tang +3
Recently, speech enhancement (SE) based on deep speech prior has attracted much attention, such as the variational auto-encoder with non-negative matrix factorization (VAE-NMF) arc…
AP20-OLR Challenge: Three Tasks and Their Baselines
Zheng Li, Miao Zhao, Qingyang Hong +5
This paper introduces the fifth oriental language recognition (OLR) challenge AP20-OLR, which intends to improve the performance of language recognition systems, along with APSIPA…