25 citations · 42 across the 4 of their papers we have counts for
7 papers · 1 filter
Accelerated Training via Incrementally Growing Neural Networks using Variance Transfer and Learning Rate Adaptation
Xin Yuan, Pedro Savarese, Michael Maire
We develop an approach to efficiently grow neural networks, within which parameterization and optimization strategies are designed by considering their effects on the training dyna…
Evaluating Machine Learning Models with NERO: Non-Equivariance Revealed on Orbits
Zhuokai Zhao, Takumi Matsuzawa, William Irvine +2
Proper evaluations are crucial for better understanding, troubleshooting, interpreting model behaviors and further improving model performance. While using scalar-based error metri…
Orthogonalized SGD and Nested Architectures for Anytime Neural Networks
Chengcheng Wan, Henry Hoffmann, Shan Lu +1
We propose a novel variant of SGD customized for training network architectures that support anytime behavior: such networks produce a series of increasingly accurate outputs over…
Winning the Lottery with Continuous Sparsification
Pedro Savarese, Hugo Silva, Michael Maire
The search for efficient, sparse deep neural network models is most prominently performed by pruning: training a dense, overparameterized network and removing parameters, usually v…
Domain-independent Dominance of Adaptive Methods
Pedro Savarese, David McAllester, Sudarshan Babu +1
From a simplified analysis of adaptive methods, we derive AvaGrad, a new optimizer which outperforms SGD on vision tasks when its adaptability is properly tuned. We observe that th…
Multigrid Neural Memory
Tri Huynh, Michael Maire, Matthew R. Walter
We introduce a novel approach to endowing neural networks with emergent, long-term, large-scale memory. Distinct from strategies that connect neural networks to external memory ban…