On Training Implicit Models
arXiv:2111.05177
Abstract
This paper focuses on training implicit models of infinite layers. Specifically, previous works employ implicit differentiation and solve the exact gradient for the backward propagation. However, is it necessary to compute such an exact but expensive gradient for training? In this work, we propose a novel gradient estimate for implicit models, named phantom gradient, that 1) forgoes the costly computation of the exact gradient; and 2) provides an update direction empirically preferable to the implicit model training. We theoretically analyze the condition under which an ascent direction of the loss landscape could be found, and provide two specific instantiations of the phantom gradient based on the damped unrolling and Neumann series. Experiments on large-scale tasks demonstrate that these lightweight phantom gradients significantly accelerate the backward passes in training implicit models by roughly 1.7 times, and even boost the performance over approaches based on the exact gradient on ImageNet.
24 pages, 4 figures, in The 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
References in corpus (13)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Improving neural networks by preventing co-adaptation of feature detectors
- Shake-Shake regularization
- Deep Equilibrium Models
- How Auto-Encoders Could Provide Credit Assignment in Deep Networks via Target Propagation
- Implicit Graph Neural Networks
- Is Attention Better Than Matrix Decomposition?
- Understanding Synthetic Gradients and Decoupled Neural Interfaces
- Implicit Feature Pyramid Network for Object Detection
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers
- Stabilizing Equilibrium Models by Jacobian Regularization
- Implicit Normalizing Flows
- Optimization Induced Equilibrium Networks