65 citations · 197 across the 9 of their papers we have counts for
3 papers
Training Transformers Together
Alexander Borzunov, Max Ryabinin, Tim Dettmers +5
The infrastructure necessary for training state-of-the-art models is becoming overly expensive, which makes training such models affordable only to large corporations and instituti…
Variable Computation in Recurrent Neural Networks
Yacine Jernite, Edouard Grave, Armand Joulin +1
Recurrent neural networks (RNNs) have been used extensively and with increasing success to model various types of sequential data. Much of this progress has been achieved through d…
Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation
Yacine Jernite, Anna Choromanska, David Sontag
We consider multi-class classification where the predictor has a hierarchical structure that allows for a very large number of labels both at train and test time. The predictive po…