activity
20162022
most citedConservative Q-Learning for Offline Reinforcement Learning

538 citations · 1.4k across the 16 of their papers we have counts for

collaborators
Showing 2017Show all

7 papers · 1 filter

cs.CL2017★ 3 cited

An online sequence-to-sequence model for noisy speech recognition

Chung-Cheng Chiu, Dieterich Lawson, Yuping Luo +4

Generative models have long been the dominant approach for speech recognition. The success of these models however relies on the use of sophisticated recipes and complicated machin…

cs.CL2017

Max-Pooling Loss Training of Long Short-Term Memory Networks for Small-Footprint Keyword Spotting

Ming Sun, Anirudh Raju, George Tucker +6

We propose a max-pooling based loss function for training Long Short-Term Memory (LSTM) networks for small-footprint keyword spotting (KWS), with low CPU, memory, and latency requi…

cs.LG2017

Filtering Variational Objectives

Chris J. Maddison, Dieterich Lawson, George Tucker +5

When used as a surrogate objective for maximum likelihood estimation in latent variable models, the evidence lower bound (ELBO) produces state-of-the-art results. Inspired by this,…

cs.AI2017

Learning Hard Alignments with Variational Inference

Dieterich Lawson, Chung-Cheng Chiu, George Tucker +3

There has recently been significant interest in hard attention models for tasks such as object recognition, visual captioning and speech recognition. Hard attention can offer benef…

cs.LG2017★ 8 cited

Particle Value Functions

Chris J. Maddison, Dieterich Lawson, George Tucker +4

The policy gradients of the expected return objective can react slowly to rare rewards. Yet, in some cases agents may wish to emphasize the low or high returns regardless of their…

cs.LG2017

REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models

George Tucker, Andriy Mnih, Chris J. Maddison +2

Learning in models with discrete latent variables is challenging due to high variance gradient estimators. Generally, approaches have relied on control variates to reduce the varia…