1 paper
Kale-ab Tessera, Sara Hooker, Benjamin Rosman
Training sparse networks to converge to the same performance as dense neural architectures has proven to be elusive. Recent work suggests that initialization is the key. However, w…