Unbiasing Truncated Backpropagation Through Time
arXiv:1705.08209
Abstract
Truncated Backpropagation Through Time (truncated BPTT) is a widespread method for learning recurrent computational graphs. Truncated BPTT keeps the computational benefits of Backpropagation Through Time (BPTT) while relieving the need for a complete backtrack through the whole data sequence at every step. However, truncation favors short-term dependencies: the gradient estimate of truncated BPTT is biased, so that it does not benefit from the convergence guarantees from stochastic gradient theory. We introduce Anticipated Reweighted Truncated Backpropagation (ARTBP), an algorithm that keeps the computational benefits of truncated BPTT, while providing unbiasedness. ARTBP works by using variable truncation lengths together with carefully chosen compensation factors in the backpropagation equation. We check the viability of ARTBP on two tasks. First, a simple synthetic task where careful balancing of temporal dependencies at different scales is needed: truncated BPTT displays unreliable performance, and in worst case scenarios, divergence, while ARTBP converges reliably. Second, on Penn Treebank character-level language modelling, ARTBP slightly outperforms truncated BPTT.
Cited by in corpus (12)
- Regularizing and Optimizing LSTM Language Models
- AdaTerm: Adaptive T-Distribution Estimated Robust Moments for Noise-Robust Stochastic Gradient Optimization
- Adaptively Truncating Backpropagation Through Time to Control Gradient Bias
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- Latent Representation in Human-Robot Interaction with Explicit Consideration of Periodic Dynamics
- Prediction of the Position of External Markers Using a Recurrent Neural Network Trained With Unbiased Online Recurrent Optimization for Safe Lung Cancer Radiotherapy
- SpikePropamine: Differentiable Plasticity in Spiking Neural Networks
- Learning Composable Energy Surrogates for PDE Order Reduction
- RIANN -- A Robust Neural Network Outperforms Attitude Estimation Filters
- Gradients are Not All You Need
- Updater-Extractor Architecture for Inductive World State Representations
- Modeling Multi-Destination Trips with Sketch-Based Model