1 paper
Qi Sun, Hexin Dong, Zewei Chen +3
Gradient-based methods for the distributed training of residual networks (ResNets) typically require a forward pass of the input data, followed by back-propagating the error gradie…