most citedImproving sequence-to-sequence speech recognition training with on-the-fly data augmentation

5 citations · 9 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CV2020

Super-Human Performance in Online Low-latency Recognition of Conversational Speech

Thai-Son Nguyen, Sebastian Stueker, Alex Waibel

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it wa…

cs.CL20201 cited

ELITR Non-Native Speech Translation at IWSLT 2020

Dominik Macháček, Jonáš Kratochvíl, Sangeet Sagar +6

This paper is an ELITR system submission for the non-native speech translation task at IWSLT 2020. We describe systems for offline ASR, real-time ASR, and our cascaded approach to…

eess.AS20201 cited

Relative Positional Encoding for Speech Recognition and Direct Translation

Ngoc-Quan Pham, Thanh-Le Ha, Tuan-Nam Nguyen +5

Transformer models are powerful sequence-to-sequence architectures that are capable of directly mapping speech inputs to transcriptions or translations. However, the mechanism for…

eess.AS2020

Low Latency ASR for Simultaneous Speech Translation

Thai Son Nguyen, Jan Niehues, Eunah Cho +6

User studies have shown that reducing the latency of our simultaneous lecture translation system should be the most important goal. We therefore have worked on several techniques f…

eess.AS2020

Toward Cross-Domain Speech Recognition with End-to-End Models

Thai-Son Nguyen, Sebastian Stüker, Alex Waibel

In the area of multi-domain speech recognition, research in the past focused on hybrid acoustic models to build cross-domain and domain-invariant speech recognition systems. In thi…

eess.AS2020

High Performance Sequence-to-Sequence Model for Streaming Speech Recognition

Thai-Son Nguyen, Ngoc-Quan Pham, Sebastian Stueker +1

Recently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the c…