255 citations · 386 across the 7 of their papers we have counts for
7 papers
Multi-Pass Transformer for Machine Translation
Peng Gao, Chiori Hori, Shijie Geng +2
In contrast with previous approaches where information flows only towards deeper layers of a stack, we consider a multi-pass transformer (MPT) architecture in which earlier layers…
Unsupervised Speaker Adaptation using Attention-based Speaker Memory for End-to-End ASR
Leda Sarı, Niko Moritz, Takaaki Hori +1
We propose an unsupervised speaker adaptation method inspired by the neural Turing machine for end-to-end (E2E) automatic speech recognition (ASR). The proposed model contains a me…
End-to-End Multi-speaker Speech Recognition with Transformer
Xuankai Chang, Wangyou Zhang, Yanmin Qian +2
Recently, fully recurrent neural network (RNN) based end-to-end models have been proven to be effective for multi-speaker speech recognition in both the single-channel and multi-ch…
Bootstrapping deep music separation from primitive auditory grouping principles
Prem Seetharaman, Gordon Wichern, Jonathan Le Roux +1
Separating an audio scene such as a cocktail party into constituent, meaningful components is a core task in computer audition. Deep networks are the state-of-the-art approach. The…
MIMO-SPEECH: End-to-End Multi-Channel Multi-Speaker Speech Recognition
Xuankai Chang, Wangyou Zhang, Yanmin Qian +2
Recently, the end-to-end approach has proven its efficacy in monaural multi-speaker speech recognition. However, high word error rates (WERs) still prevent these systems from being…
Full-Capacity Unitary Recurrent Neural Networks
Scott Wisdom, Thomas Powers, John R. Hershey +2
Recurrent neural networks are powerful models for processing sequential data, but they are generally plagued by vanishing and exploding gradient problems. Unitary recurrent neural…