12 citations · 32 across the 7 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2022★ 8 cited
Survey on Large Scale Neural Network Training
Julia Gusak, Daria Cherniuk, Alena Shilova +8
Modern Deep Neural Networks (DNNs) require significant memory to store weight, activations, and other intermediate tensors during training. Hence, many models do not fit one GPU de…
cs.LG2022★ 5 cited
Few-Bit Backward: Quantized Gradients of Activation Functions for Memory Footprint Reduction
Georgii Novikov, Daniel Bershatsky, Julia Gusak +3
Memory footprint is one of the main limiting factors for large neural network training. In backpropagation, one needs to store the input to each operation in the computational grap…