6 citations · 10 across the 2 of their papers we have counts for
2 papers
cs.CL2022★ 6 cited
What Language Model to Train if You Have One Million GPU Hours?
Teven Le Scao, Thomas Wang, Daniel Hesslow +16
The crystallization of modeling methods around the Transformer architecture has been a boon for practitioners. Simple, well-motivated architectural variations can transfer across t…
cs.LG2021★ 4 cited
Distributed Deep Learning in Open Collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin +13
Modern deep learning applications require increasingly more compute to train state-of-the-art models. To address this demand, large corporations and institutions use dedicated High…