1 paper
Lu Xia, Michiel E. Hochstenbach, Stefano Massei
When training neural networks with low-precision computation, rounding errors often cause stagnation or are detrimental to the convergence of the optimizers; in this paper we study…