3 papers
cs.SD2026
ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning
Khanh Le, Kiet Anh Hoang, Bao Nguyen +5
We present ViP-VL, an efficient Vietnamese Self-supervised speech Pretraining model leveraging Vector-quantization Learning. To bridge the gap between high-resolution audio and eff…
cs.SD2025
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
Khanh Le, Tuan Vu Ho, Dung Tran +1
RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor t…
cs.SD2025
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
Khanh Le, Tuan Vu Ho, Dung Tran +1
Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hour…