3 papers
cs.CL2025
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
Dancheng Liu, Amir Nassereldine, Chenhui Xu +1
Whisper's robust performance in automatic speech recognition (ASR) is often attributed to its massive 680k-hour training set, an impractical scale for most researchers. In this wor…
cs.LG2024
NVCiM-PT: An NVCiM-assisted Prompt Tuning Framework for Edge LLMs
Ruiyang Qin, Pengyu Ren, Zheyu Yan +7
Large Language Models (LLMs) deployed on edge devices, known as edge LLMs, need to continuously fine-tune their model parameters from user-generated data under limited resource con…
cs.CL2024
FASA: a Flexible and Automatic Speech Aligner for Extracting High-quality Aligned Children Speech Data
Dancheng Liu, Jinjun Xiong
Automatic Speech Recognition (ASR) for adults' speeches has made significant progress by employing deep neural network (DNN) models recently, but improvement in children's speech i…