2.1k citations · 3.3k across the 18 of their papers we have counts for
7 papers · 1 filter
Meta-Learning Fast Weight Language Models
Kevin Clark, Kelvin Guu, Ming-Wei Chang +3
Dynamic evaluation of language models (LMs) adapts model parameters at test time using gradient information from previous tokens and substantially improves LM performance. However,…
SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network
William Chan, Daniel Park, Chris Lee +3
We present SpeechStew, a speech recognition model that is trained on a combination of various publicly available speech recognition datasets: AMI, Broadcast News, Common Voice, Lib…
Dynamic Programming Encoding for Subword Segmentation in Neural Machine Translation
Xuanli He, Gholamreza Haffari, Mohammad Norouzi
This paper introduces Dynamic Programming Encoding (DPE), a new segmentation algorithm for tokenizing sentences into subword units. We view the subword segmentation of output sente…
Non-Autoregressive Machine Translation with Latent Alignments
Chitwan Saharia, William Chan, Saurabh Saxena +1
This paper presents two strong methods, CTC and Imputer, for non-autoregressive machine translation that model latent alignments with dynamic programming. We revisit CTC for machin…
Sequence to Sequence Mixture Model for Diverse Machine Translation
Xuanli He, Gholamreza Haffari, Mohammad Norouzi
Sequence to sequence (SEQ2SEQ) models often lack diversity in their generated translations. This can be attributed to the limitation of SEQ2SEQ models in capturing lexical and synt…
Embedding Text in Hyperbolic Spaces
Bhuwan Dhingra, Christopher J. Shallue, Mohammad Norouzi +2
Natural language text exhibits hierarchical structure in a variety of respects. Ideally, we could incorporate our prior knowledge of this hierarchical structure into unsupervised l…