538 citations · 1.4k across the 16 of their papers we have counts for
7 papers · 1 filter
An online sequence-to-sequence model for noisy speech recognition
Chung-Cheng Chiu, Dieterich Lawson, Yuping Luo +4
Generative models have long been the dominant approach for speech recognition. The success of these models however relies on the use of sophisticated recipes and complicated machin…
Max-Pooling Loss Training of Long Short-Term Memory Networks for Small-Footprint Keyword Spotting
Ming Sun, Anirudh Raju, George Tucker +6
We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requi…
Filtering Variational Objectives
Chris J. Maddison, Dieterich Lawson, George Tucker +5
When used as a surrogate objective for maximum likelihood estimation in latent variable models, the evidence lower bound (ELBO) produces state-of-the-art results. Inspired by this,…
Learning Hard Alignments with Variational Inference
Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3
There has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention can offer benef…
Particle Value Functions
Chris J. Maddison, Dieterich Lawson, George Tucker +4
The policy gradients of the expected return objective can react slowly to rare rewards. Yet, in some cases agents may wish to emphasize the low or high returns regardless of their…
REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models
George Tucker, Andriy Mnih, Chris J. Maddison +2
Learning in models with discrete latent variables is challenging due to high variance gradient estimators. Generally, approaches have relied on control variates to reduce the varia…