24 citations · 52 across the 7 of their papers we have counts for
7 papers
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
Sneha Kudugunta, Isaac Caswell, Biao Zhang +8
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-a…
Cross-Lingual Supervision improves Large Language Models Pre-training
Andrea Schioppa, Xavier Garcia, Orhan Firat
The recent rapid progress in pre-training Large Language Models has relied on using self-supervised language modeling objectives like next token prediction or span corruption. On t…
UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
Hyung Won Chung, Noah Constant, Xavier Garcia +4
Pretrained multilingual large language models have typically used heuristic temperature-based sampling to balance between different languages. However previous work has not systema…
Scaling Laws for Multilingual Neural Machine Translation
Patrick Fernandes, Behrooz Ghorbani, Xavier Garcia +2
In this work, we provide a large-scale empirical study of the scaling properties of multilingual neural machine translation models. We examine how increases in the model size affec…
The unreasonable effectiveness of few-shot learning for machine translation
Xavier Garcia, Yamini Bansal, Colin Cherry +5
We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples…
Measuring The Impact Of Programming Language Distribution
Gabriel Orlanski, Kefan Xiao, Xavier Garcia +6
Current benchmarks for evaluating neural code models focus on only a small subset of programming languages, excluding many popular languages such as Go or Rust. To ameliorate this…