6 citations · 6 across the 1 of their papers we have counts for
3 papers
cs.LG2023★ 1 cited
A Computationally Efficient Sparsified Online Newton Method
Fnu Devvrit, Sai Surya Duvvuri, Rohan Anil +3
Second-order methods hold significant promise for enhancing the convergence of deep neural network training; however, their large memory and computational demands have limited thei…
cs.LG2023
Heterogeneous Federated Learning Using Knowledge Codistillation
Jared Lichtarge, Ehsan Amid, Shankar Kumar +3
Federated Averaging, and many federated learning algorithm variants which build upon it, have a limitation: all clients must share the same model architecture. This results in unus…
cs.CL2022★ 6 cited
N-Grammer: Augmenting Transformers with latent n-grams
Aurko Roy, Rohan Anil, Guangda Lai +13
Transformer models have recently emerged as one of the foundational models in natural language processing, and as a byproduct, there is significant recent interest and investment i…