1 paper
Brian Chmiel, Liad Ben-Uri, Moran Shkolnik +3
While training can mostly be accelerated by reducing the time needed to propagate neural gradients back throughout the model, most previous works focus on the quantization/pruning…