3 papers
cs.LG2026
Revisiting ASR Error Correction with Specialized Models
Zijin Gu, Tatiana Likhomanenko, He Bai +3
Language models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models…
cs.CL2025
dMel: Speech Tokenization made Simple
Richard He Bai, Tatiana Likhomanenko, Ruixiang Zhang +3
Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have inv…
eess.AS2025
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
Li-Wei Chen, Takuya Higuchi, He Bai +6
Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use…