3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.LG2022★ 3 cited
Exploring Low Rank Training of Deep Neural Networks
Siddhartha Rao Kamalakara, Acyr Locatelli, Bharat Venkitesh +3
Training deep neural networks in low rank, i.e. with factorised layers, is of particular interest to the community: it offers efficiency over unfactorised training in terms of both…
cs.LG2022★ 1 cited
Scalable Training of Language Models using JAX pjit and TPUv4
Joanna Yoo, Kuba Perlin, Siddhartha Rao Kamalakara +1
Modern large language models require distributed training strategies due to their size. The challenges of efficiently and robustly training them are met with rapid developments on…