3 citations · 3 across the 3 of their papers we have counts for
5 papers
Autoregressive Knowledge Distillation through Imitation Learning
Alexander Lin, Jeremy Wohlwend, Howard Chen +1
The performance of autoregressive models on natural language generation tasks has dramatically improved due to the adoption of deep, self-attentive architectures. However, these ga…
ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech Recognition
Jing Pan, Joshua Shapiro, Jeremy Wohlwend +3
In this paper we present state-of-the-art (SOTA) performance on the LibriSpeech corpus with two novel neural network architectures, a multistream CNN for acoustic modeling and a se…
Metric Learning for Dynamic Text Classification
Jeremy Wohlwend, Ethan R. Elenberg, Samuel Altschul +2
Traditional text classifiers are limited to predicting over a fixed set of labels. However, in many real-world applications the label set is frequently changing. For example, in in…
Structured Pruning of Large Language Models
Ziheng Wang, Jeremy Wohlwend, Tao Lei
Large language models have recently achieved state of the art performance across a wide variety of natural language tasks. Meanwhile, the size of these models and their latency hav…
Building a Production Model for Retrieval-Based Chatbots
Kyle Swanson, Lili Yu, Christopher Fox +2
Response suggestion is an important task for building human-computer conversation systems. Recent approaches to conversation modeling have introduced new model architectures with i…