48 citations · 49 across the 2 of their papers we have counts for
6 papers
Scaling Up Models and Data with and
Adam Roberts, Hyung Won Chung, Anselm Levskaya +40
Recent neural network-based language models have benefited greatly from scaling up the size of training datasets and the number of parameters in the models themselves. Scaling can…
Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning
Nan Ding, Xi Chen, Tomer Levinboim +2
Despite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain perform…
TeaForN: Teacher-Forcing with N-grams
Sebastian Goodman, Nan Ding, Radu Soricut
Sequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps. Our proposed method, Teacher-For…
Multi-Image Summarization: Textual Summary from a Set of Cohesive Images
Nicholas Trieu, Sebastian Goodman, Pradyumna Narayana +2
Multi-sentence summarization is a well studied problem in NLP, while generating image descriptions for a single image is a well studied problem in Computer Vision. However, for app…
Multi-stage Pretraining for Abstractive Summarization
Sebastian Goodman, Zhenzhong Lan, Radu Soricut
Neural models for abstractive summarization tend to achieve the best performance in the presence of highly specialized, summarization specific modeling add-ons such as pointer-gene…
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman +3
Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks. However, at some point further model increases be…