2 papers
cs.SD2025
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
Khanh Le, Tuan Vu Ho, Dung Tran +1
RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor t…
cs.SD2025
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
Khanh Le, Tuan Vu Ho, Dung Tran +1
Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hour…