58 citations · 123 across the 4 of their papers we have counts for
12 papers
Multilingual Machine Translation with Open Large Language Models at Practical Scale: An Empirical Study
Menglong Cui, Pengzhi Gao, Wei Liu +2
Large language models (LLMs) have shown continuously improving multilingual capabilities, and even small-scale open-source models have demonstrated rapid performance enhancement. I…
Scaling Up Models and Data with and
Adam Roberts, Hyung Won Chung, Anselm Levskaya +40
Recent neural network-based language models have benefited greatly from scaling up the size of training datasets and the number of parameters in the models themselves. Scaling can…
How to decay your learning rate
Aitor Lewkowycz
Complex learning rate schedules have become an integral part of deep learning. We find empirically that common fine-tuned schedules decay the learning rate after the weight norm bo…
Gravitational path integral from the deformation
Alexandre Belin, Aitor Lewkowycz, Gabor Sarosi
We study a deformation of large conformal field theories, a higher dimensional generalization of the deformation. The deformed partition function satisfies a fl…
On the training dynamics of deep networks with regularization
Aitor Lewkowycz, Guy Gur-Ari
We study the role of regularization in deep learning, and uncover simple relations between the performance of the model, the coefficient, the learning rate, and the num…
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer +2
The choice of initial learning rate can have a profound effect on the performance of deep networks. We present a class of neural networks with solvable training dynamics, and confi…