58 citations · 133 across the 6 of their papers we have counts for
4 papers · 1 filter
Pre-trained Summarization Distillation
Sam Shleifer, Alexander M. Rush
Recent state-of-the-art approaches to summarization utilize large pre-trained Transformer models. Distilling these models to smaller student models has become critically important…
Classification as Decoder: Trading Flexibility for Control in Medical Dialogue
Sam Shleifer, Manish Chablani, Anitha Kannan +2
Generative seq2seq dialogue systems are trained to predict the next word in dialogues that have already occurred. They can learn from large unlabeled conversation datasets, build a…
Classification As Decoder: Trading Flexibility For Control In Neural Dialogue
Sam Shleifer, Manish Chablani, Namit Katariya +2
Generative seq2seq dialogue systems are trained to predict the next word in dialogues that have already occurred. They can learn from large unlabeled conversation datasets, build a…
Low Resource Text Classification with ULMFit and Backtranslation
Sam Shleifer
In computer vision, virtually every state-of-the-art deep learning system is trained with data augmentation. In text classification, however, data augmentation is less widely pract…