most citedPre-trained Summarization Distillation

58 citations · 133 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL202058 cited

Pre-trained Summarization Distillation

Sam Shleifer, Alexander M. Rush

Recent state-of-the-art approaches to summarization utilize large pre-trained Transformer models. Distilling these models to smaller student models has become critically important…

eess.SP201917 cited

Incrementally Improving Graph WaveNet Performance on Traffic Prediction

Sam Shleifer, Clara McCreery, Vamsi Chitters

We present a series of modifications which improve upon Graph WaveNet's previously state-of-the-art performance on the METR-LA traffic prediction task. The goal of this task is to…

cs.CL2019

Classification as Decoder: Trading Flexibility for Control in Medical Dialogue

Sam Shleifer, Manish Chablani, Anitha Kannan +2

Generative seq2seq dialogue systems are trained to predict the next word in dialogues that have already occurred. They can learn from large unlabeled conversation datasets, build a…

cs.CL2019

Classification As Decoder: Trading Flexibility For Control In Neural Dialogue

Sam Shleifer, Manish Chablani, Namit Katariya +2

Generative seq2seq dialogue systems are trained to predict the next word in dialogues that have already occurred. They can learn from large unlabeled conversation datasets, build a…

cs.LG201915 cited

Using Small Proxy Datasets to Accelerate Hyperparameter Search

Sam Shleifer, Eric Prokop

One of the biggest bottlenecks in a machine learning workflow is waiting for models to train. Depending on the available computing resources, it can take days to weeks to train a n…

cs.CL201943 cited

Low Resource Text Classification with ULMFit and Backtranslation

Sam Shleifer

In computer vision, virtually every state-of-the-art deep learning system is trained with data augmentation. In text classification, however, data augmentation is less widely pract…