1 paper
Jose Javier Gonzalez Ortiz, Abhay Gupta, Christopher Rinard +1
Standard mixed-precision training of neural networks requires many bytes of accelerator memory for each model parameter. These bytes reflect not just the parameter itself, but also…