5 citations · 6 across the 2 of their papers we have counts for
4 papers
Relative Positional Encoding for Speech Recognition and Direct Translation
Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen +5
Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for…
Low Latency ASR for Simultaneous Speech Translation
Thai Son Nguyen, Jan Niehues, Eunah Cho +6
User studies have shown that reducing the latency of our simultaneous lecture translation system should be the most important goal. We therefore have worked on several techniques f…
High Performance Sequence-to-Sequence Model for Streaming Speech Recognition
Thai-Son Nguyen, Ngoc-Quan Pham, Sebastian Stueker +1
Recently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the c…
Improving sequence-to-sequence speech recognition training with on-the-fly data augmentation
Thai-Son Nguyen, Sebastian Stueker, Jan Niehues +1
Sequence-to-Sequence (S2S) models recently started to show state-of-the-art performance for automatic speech recognition (ASR). With these large and deep models overfitting remains…