31 citations · 62 across the 8 of their papers we have counts for
4 papers · 1 filter
Pretraining by Backtranslation for End-to-end ASR in Low-Resource Settings
Matthew Wiesner, Adithya Renduchintala, Shinji Watanabe +3
We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because t…
Improving End-to-end Speech Recognition with Pronunciation-assisted Sub-word Modeling
Hainan Xu, Shuoyang Ding, Shinji Watanabe
Most end-to-end speech recognition systems model text directly as a sequence of characters or sub-words. Current approaches to sub-word extraction only consider character sequence…
How Do Source-side Monolingual Word Embeddings Impact Neural Machine Translation?
Shuoyang Ding, Kevin Duh
Using pre-trained word embeddings as input layer is a common practice in many natural language processing (NLP) tasks, but it is largely neglected for neural machine translation (N…
Multi-Modal Data Augmentation for End-to-End ASR
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner +1
We present a new end-to-end architecture for automatic speech recognition (ASR) that can be trained using \emph{symbolic} input in addition to the traditional acoustic input. This…