5 citations · 17 across the 10 of their papers we have counts for
Showing 2020Show all
3 papers · 1 filter
cs.CL2020
Autoregressive Knowledge Distillation through Imitation Learning
Alexander Lin, Jeremy Wohlwend, Howard Chen +1
The performance of autoregressive models on natural language generation tasks has dramatically improved due to the adoption of deep, self-attentive architectures. However, these ga…
cs.LG2020★ 2 cited
Rationalizing Text Matching: Learning Sparse Alignments via Optimal Transport
Kyle Swanson, Lili Yu, Tao Lei
Selecting input features of top relevance has become a popular method for building self-explaining models. In this work, we extend this selective rationalization approach to text m…
eess.AS2020★ 3 cited
ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition
Jing Pan, Joshua Shapiro, Jeremy Wohlwend +3
In this paper we present state-of-the-art (SOTA) performance on the LibriSpeech corpus with two novel neural network architectures, a multistream CNN for acoustic modeling and a se…