580 citations · 1.8k across the 13 of their papers we have counts for
25 papers
Multi-step Planning for Automated Hyperparameter Optimization with OptFormer
Lucio M. Dery, Abram L. Friesen, Nando De Freitas +2
As machine learning permeates more industries and models become more expensive and time consuming to train, the need for efficient automated hyperparameter optimization (HPO) has n…
The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary +7
One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks eithe…
Few-shot Sequence Learning with Transformers
Lajanugen Logeswaran, Ann Lee, Myle Ott +3
Few-shot algorithms aim at learning new tasks provided only a handful of training examples. In this work we investigate few-shot learning in the setting where the data points are s…
Efficient Continual Learning with Modular Networks and Task-Driven Priors
Tom Veniat, Ludovic Denoyer, Marc'Aurelio Ranzato
Existing literature in Continual Learning (CL) has focused on overcoming catastrophic forgetting, the inability of the learner to recall how to perform tasks observed in the past.…
Multi-scale Transformer Language Models
Sandeep Subramanian, Ronan Collobert, Marc'Aurelio Ranzato +1
We investigate multi-scale transformer language models that learn representations of text at multiple scales, and present three different architectures that have an inductive bias…
Residual Energy-Based Models for Text
Anton Bakhtin, Yuntian Deng, Sam Gross +3
Current large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations o…