4 papers
Dynamic Context-Aware Streaming Pretrained Language Model For Inverse Text Normalization
Luong Ho, Khanh Le, Vinh Pham +3
Inverse Text Normalization (ITN) is crucial for converting spoken Automatic Speech Recognition (ASR) outputs into well-formatted written text, enhancing both readability and usabil…
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
Khanh Le, Duc Chau
Chunk-based inference stands out as a popular approach in developing real-time streaming speech recognition, valued for its simplicity and efficiency. However, because it restricts…
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
Khanh Le, Tuan Vu Ho, Dung Tran +1
RNN-Transducer (RNN-T) is a widely adopted architecture in speech recognition, integrating acoustic and language modeling in an end-to-end framework. However, the RNN-T predictor t…
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
Khanh Le, Tuan Vu Ho, Dung Tran +1
Deploying ASR models at an industrial scale poses significant challenges in hardware resource management, especially for long-form transcription tasks where audio may last for hour…