113 citations · 130 across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 6 cited
UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
Hyung Won Chung, Noah Constant, Xavier Garcia +4
Pretrained multilingual large language models have typically used heuristic temperature-based sampling to balance between different languages. However previous work has not systema…
cs.AI2023★ 113 cited
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
Shayne Longpre, Le Hou, Tu Vu +8
We study the design decisions of publicly available instruction tuning methods, and break down the development of Flan 2022 (Chung et al., 2022). Through careful ablation studies o…
cs.LG2022★ 11 cited
Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?
Yi Tay, Mostafa Dehghani, Samira Abnar +7
There have been a lot of interest in the scaling properties of Transformer models. However, not much has been done on the front of investigating the effect of scaling properties of…