2 citations · 2 across the 1 of their papers we have counts for
1 paper
Joya Chen, Kai Xu, Yuhui Wang +2
A standard hardware bottleneck when training deep neural networks is GPU memory. The bulk of memory is occupied by caching intermediate tensors for gradient computation in the back…