205 citations · 577 across the 18 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.LG2023★ 8 cited
No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language Models
Jean Kaddour, Oscar Key, Piotr Nawrot +2
The computation necessary for training Transformer-based language models has skyrocketed in recent years. This trend has motivated research on efficient training algorithms designe…
cs.LG2023
DAG Learning on the Permutahedron
Valentina Zantedeschi, Luca Franceschi, Jean Kaddour +2
We propose a continuous optimization framework for discovering a latent directed acyclic graph (DAG) from observational data. Our approach optimizes over the polytope of permutatio…