1 paper
Peter Izsak, Moshe Berchansky, Omer Levy
While large language models a la BERT are used ubiquitously in NLP, pretraining them is considered a luxury that only a few well-funded industry labs can afford. How can one train…