1 paper
Neel Mishra, Kushagara Trivedi, Pawan Kumar
Distributed training of large neural networks is bottlenecked by full-precision gradient communication and by coordinatewise optimizers that ignore the matrix structure of weight t…