Publications (34)
Real or Fake? Learning to Discriminate Machine from Human Generated Text
Anton Bakhtin, Sam Gross, Myle Ott +3
Energy-based models (EBMs), a.k.a. un-normalized models, have had recent successes in continuous spaces. However, they have not been successfully applied to model text sequences. W…
The FLoRes Evaluation Datasets for Low-Resource Machine Translation: Nepali-English and Sinhala-English
Francisco Guzmán, Peng-Jen Chen, Myle Ott +5
For machine translation, a vast majority of language pairs in the world are considered low-resource because they have little parallel data available. Besides the technical challeng…
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal +9
Building open-domain chatbots is a challenging area for machine learning research. While prior work has shown that scaling neural models in the number of parameters and the size of…
Phrase-Based & Neural Unsupervised Machine Translation
Guillaume Lample, Myle Ott, Alexis Conneau +2
Machine translation systems achieve near human-level performance on some languages, yet their effectiveness strongly relies on the availability of large amounts of parallel sentenc…
Scaling Neural Machine Translation
Myle Ott, Sergey Edunov, David Grangier +1
Sequence to sequence learning models still require several days to reach state of the art performance on large benchmark datasets using a single machine. This paper shows that redu…
Few-shot Learning with Multilingual Language Models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe +18
Large-scale generative language models such as GPT-3 are competitive few-shot learners. While these models are known to be able to jointly represent many different languages, their…
Estimating the Prevalence of Deception in Online Review Communities
Myle Ott, Claire Cardie, Jeff Hancock
Consumers' purchase decisions are increasingly influenced by user-generated online reviews. Accordingly, there has been growing concern about the potential for posting "deceptive o…
fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Myle Ott, Sergey Edunov, Alexei Baevski +5
fairseq is an open-source sequence modeling toolkit that allows researchers and developers to train custom models for translation, summarization, language modeling, and other text…
The Source-Target Domain Mismatch Problem in Machine Translation
Jiajun Shen, Peng-Jen Chen, Matt Le +5
While we live in an increasingly interconnected world, different places still exhibit strikingly different cultures and many events we experience in our every day life pertain only…
Analyzing the Forgetting Problem in the Pretrain-Finetuning of Dialogue Response Models
Tianxing He, Jun Liu, Kyunghyun Cho +4
In this work, we study how the finetuning stage in the pretrain-finetune framework changes the behavior of a pretrained neural language generator. We focus on the transformer encod…
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal +16
Large language models, which are often trained for hundreds of thousands of compute days, have shown remarkable capabilities for zero- and few-shot learning. Given their computatio…
NormFormer: Improved Transformer Pretraining with Extra Normalization
Sam Shleifer, Jason Weston, Myle Ott
During pretraining, the Pre-LayerNorm transformer suffers from a gradient magnitude mismatch: gradients at early layers are much larger than at later layers. These issues can be al…
Few-shot Sequence Learning with Transformers
Lajanugen Logeswaran, Ann Lee, Myle Ott +3
Few-shot algorithms aim at learning new tasks provided only a handful of training examples. In this work we investigate few-shot learning in the setting where the data points are s…
Facebook FAIR's WMT19 News Translation Task Submission
Nathan Ng, Kyra Yee, Alexei Baevski +3
This paper describes Facebook FAIR's submission to the WMT19 shared news translation task. We participate in two language pairs and four language directions, English <-> German and…
Efficient Large Scale Language Modeling with Mixtures of Experts
Mikel Artetxe, Shruti Bhosale, Naman Goyal +21
Mixture of Experts layers (MoEs) enable efficient scaling of language models through conditional computation. This paper presents a detailed empirical study of how autoregressive M…
How Decoding Strategies Affect the Verifiability of Generated Text
Luca Massarelli, Fabio Petroni, Aleksandra Piktus +5
Recent progress in pre-trained language models led to systems that are able to generate text of an increasingly high quality. While several works have investigated the fluency and…
Finding Deceptive Opinion Spam by Any Stretch of the Imagination
Myle Ott, Yejin Choi, Claire Cardie +1
Consumers increasingly rate, review and research products online. Consequently, websites containing consumer reviews are becoming targets of opinion spam. While recent work has foc…
Efficient Language Modeling with Sparse all-MLP
Ping Yu, Mikel Artetxe, Myle Ott +4
All-MLP architectures have attracted increasing interest as an alternative to attention-based models. In NLP, recent work like gMLP shows that all-MLPs can match Transformers in la…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
On The Evaluation of Machine Translation Systems Trained With Back-Translation
Sergey Edunov, Myle Ott, Marc'Aurelio Ranzato +1
Back-translation is a widely used data augmentation technique which leverages target monolingual data. However, its effectiveness has been challenged since automatic metrics such a…
Analyzing Uncertainty in Neural Machine Translation
Myle Ott, Michael Auli, David Grangier +1
Machine translation is a popular test bed for research in neural sequence-to-sequence models but despite much recent research, there is still a lack of understanding of these model…
Larger-Scale Transformers for Multilingual Masked Language Modeling
Naman Goyal, Jingfei Du, Myle Ott +2
Recent work has demonstrated the effectiveness of cross-lingual language model pretraining for cross-lingual understanding. In this study, we present the results of two larger mult…
Residual Energy-Based Models for Text
Anton Bakhtin, Yuntian Deng, Sam Gross +3
Current large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations o…
On Anytime Learning at Macroscale
Lucas Caccia, Jing Xu, Myle Ott +2
In many practical applications of machine learning data arrives sequentially over time in large chunks. Practitioners have then to decide how to allocate their computational budget…
Sustainable AI: Environmental Implications, Challenges and Opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta +22
This paper explores the environmental impact of the super-linear growth trends for AI from a holistic perspective, spanning Data, Algorithms, and System Hardware. We characterize t…
Facebook AI's WAT19 Myanmar-English Translation Task Submission
Peng-Jen Chen, Jiajun Shen, Matt Le +5
This paper describes Facebook AI's submission to the WAT 2019 Myanmar-English translation task. Our baseline systems are BPE-based transformer models. We explore methods to leverag…
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal +7
Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often…
General Purpose Text Embeddings from Pre-trained Language Models for Scalable Inference
Jingfei Du, Myle Ott, Haoran Li +2
The state of the art on many NLP tasks is currently achieved by large pre-trained language models, which require a considerable amount of computation. We explore a setting where ma…
Classical Structured Prediction Losses for Sequence to Sequence Learning
Sergey Edunov, Myle Ott, Michael Auli +2
There has been much recent work on training neural attention models at the sequence-level using either reinforcement learning-style methods or by optimizing the beam. In this paper…
Residual Energy-Based Models for Text Generation
Yuntian Deng, Anton Bakhtin, Myle Ott +2
Text generation is ubiquitous in many NLP tasks, from summarization, to dialogue and machine translation. The dominant parametric approach is based on locally normalized models whi…
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Yanli Zhao, Andrew Gu, Rohan Varma +15
It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of…
Understanding Back-Translation at Scale
Sergey Edunov, Myle Ott, Michael Auli +1
An effective method to improve neural machine translation with monolingual data is to augment the parallel training corpus with back-translations of target language sentences. This…
Unsupervised Cross-lingual Representation Learning at Scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal +7
This paper shows that pretraining multilingual language models at scale leads to significant performance gains for a wide range of cross-lingual transfer tasks. We train a Transfor…
Mixture Models for Diverse Machine Translation: Tricks of the Trade
Tianxiao Shen, Myle Ott, Michael Auli +1
Mixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature. Surprisingly, these models hav…