Adaptive Computation Time for Recurrent Neural Networks
arXiv:1603.08983
Abstract
This paper introduces Adaptive Computation Time (ACT), an algorithm that allows recurrent neural networks to learn how many computational steps to take between receiving an input and emitting an output. ACT requires minimal changes to the network architecture, is deterministic and differentiable, and does not add any noise to the parameter gradients. Experimental results are provided for four synthetic problems: determining the parity of binary vectors, applying binary logic operations, adding integers, and sorting real numbers. Overall, performance is dramatically improved by the use of ACT, which successfully adapts the number of computational steps to the requirements of the problem. We also present character-level language modelling results on the Hutter prize Wikipedia dataset. In this case ACT does not yield large gains in performance; however it does provide intriguing insight into the structure of the data, with more computation allocated to harder-to-predict transitions, such as spaces between words and ends of sentences. This suggests that ACT or other adaptive computation methods could provide a generic method for inferring segment boundaries in sequence data.
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- DRAW: A Recurrent Neural Network For Image Generation
- Attend, Infer, Repeat: Fast Scene Understanding with Generative Models
- Neural Programmer-Interpreters
- Order Matters: Sequence to sequence for sets
- Conditional Computation in Neural Networks for faster models
- Multi-column Deep Neural Networks for Image Classification
- Deep Sequential Neural Network
- Self-Delimiting Neural Networks
Cited by in corpus (131)
- Language Models are Few-Shot Learners
- Pre-trained Models for Natural Language Processing: A Survey
- Neural Ordinary Differential Equations
- Principal Neighbourhood Aggregation for Graph Nets
- Attend, Infer, Repeat: Fast Scene Understanding with Generative Models
- Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
- Bringing AI To Edge: From Deep Learning's Perspective
- The Predictron: End-To-End Learning and Planning
- A Practical Survey on Faster and Lighter Transformers
- Adaptive Propagation Graph Convolutional Network
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Analysing Mathematical Reasoning Abilities of Neural Models
- Review Networks for Caption Generation
- Learning model-based planning from scratch
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- Depth-Adaptive Transformer
- Sequence-to-sequence neural network models for transliteration
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory
- Pruning and Quantization for Deep Neural Network Acceleration: A Survey
- BERT Loses Patience: Fast and Robust Inference with Early Exit
- Fast-Slow Recurrent Neural Networks
- Metacontrol for Adaptive Imagination-Based Optimization
- Dynamic Neural Networks: A Survey
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Simple, Distributed, and Accelerated Probabilistic Programming
- NAIS-Net: Stable Deep Networks from Non-Autonomous Differential Equations
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- Adversarial Robustness for Code
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Dynamically Sacrificing Accuracy for Reduced Computation: Cascaded Inference Based on Softmax Confidence
- IamNN: Iterative and Adaptive Mobile Neural Network for Efficient Image Classification
- Learning to Segment Inputs for NMT Favors Character-Level Processing
- Attention over Parameters for Dialogue Systems
- Focused Hierarchical RNNs for Conditional Sequence Processing
- DDPG-Driven Deep-Unfolding with Adaptive Depth for Channel Estimation with Sparse Bayesian Learning
- Spatially Adaptive Computation Time for Residual Networks
- Lessons on Parameter Sharing across Layers in Transformers
- Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
- Deep Image Harmonization in Dual Color Spaces
- PonderNet: Learning to Ponder
- AdaFuse: Adaptive Temporal Fusion Network for Efficient Action Recognition
- Attending to Mathematical Language with Transformers
- Improved Techniques for Training Adaptive Deep Networks
- Dynamic Computational Time for Visual Attention
- How hard is to distinguish graphs with graph neural networks?
- Efficient Transformers with Dynamic Token Pooling
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- AdaFrame: Adaptive Frame Selection for Fast Video Recognition
- CSPN++: Learning Context and Resource Aware Convolutional Spatial Propagation Networks for Depth Completion
- Early Exiting with Ensemble Internal Classifiers
- Learning to Remember More with Less Memorization
- Deep Amortized Clustering
- Deep Learning for Embodied Vision Navigation: A Survey
- Improving Differentiable Neural Computers Through Memory Masking, De-allocation, and Link Distribution Sharpness Control
- Controlling Computation versus Quality for Neural Sequence Models
- Extensions and Limitations of the Neural GPU
- AR-Net: Adaptive Frame Resolution for Efficient Action Recognition
- EREBA: Black-box Energy Testing of Adaptive Neural Networks
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- Variable-rate discrete representation learning
- Scalable Transformers for Neural Machine Translation
- InfoCNF: An Efficient Conditional Continuous Normalizing Flow with Adaptive Solvers
- Fighting Gradients with Gradients: Dynamic Defenses against Adversarial Attacks
- Dynamic Compositional Graph Convolutional Network for Efficient Composite Human Motion Prediction
- Computing a human-like reaction time metric from stable recurrent vision models
- Learning to Reason With Adaptive Computation
- MEMO: A Deep Network for Flexible Combination of Episodic Memories
- RomeBERT: Robust Training of Multi-Exit BERT
- Set-to-Sequence Methods in Machine Learning: a Review
- Auto-Vectorizing TensorFlow Graphs: Jacobians, Auto-Batching And Beyond
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI
- Universal Transforming Geometric Network
- Surprisal-Triggered Conditional Computation with Neural Networks
- Comparing Fixed and Adaptive Computation Time for Recurrent Neural Networks
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
- Cell-aware Stacked LSTMs for Modeling Sentences
- VA-RED: Video Adaptive Redundancy Reduction
- Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers
- Long-Distance Gesture Recognition using Dynamic Neural Networks
- On the Transformer Growth for Progressive BERT Training
- Energy-efficient Amortized Inference with Cascaded Deep Classifiers
- Early Improving Recurrent Elastic Highway Network
- Anytime Prediction as a Model of Human Reaction Time
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- Iterative Amortized Policy Optimization
- Towards Scale-Invariant Graph-related Problem Solving by Iterative Homogeneous Graph Neural Networks
- State-Denoised Recurrent Neural Networks
- Can You Learn an Algorithm? Generalizing from Easy to Hard Problems with Recurrent Networks
- CoDiNet: Path Distribution Modeling with Consistency and Diversity for Dynamic Routing
- CascadeBERT: Accelerating Inference of Pre-trained Language Models via Calibrated Complete Models Cascade
- MuFuRU: The Multi-Function Recurrent Unit
- Latent Universal Task-Specific BERT
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
- Spatiotemporal Adaptive Neural Network for Long-term Forecasting of Financial Time Series
- Adversarial Subword Regularization for Robust Neural Machine Translation
- Post-Train Adaptive U-Net for Image Segmentation
- Smoothness Sensor: Adaptive Smoothness-Transition Graph Convolutions for Attributed Graph Clustering
- Limitations of Autoregressive Models and Their Alternatives
- Multi-Zone Unit for Recurrent Neural Networks
- Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference Loss
- Efficiently applying attention to sequential data with the Recurrent Discounted Attention unit
- Tackling real noisy reverberant meetings with all-neural source separation, counting, and diarization system
- MOOD: Multi-level Out-of-distribution Detection
- Thought Flow Nets: From Single Predictions to Trains of Model Thought
- End-to-end Speech Recognition with Adaptive Computation Steps
- Memory and attention in deep learning
- Depth-Adaptive Graph Recurrent Network for Text Classification
- Spike-inspired Rank Coding for Fast and Accurate Recurrent Neural Networks
- Deep Episodic Value Iteration for Model-based Meta-Reinforcement Learning
- When in Doubt, Summon the Titans: Efficient Inference with Large Models
- Leveraging Recursive Gumbel-Max Trick for Approximate Inference in Combinatorial Spaces
- Channel selection using Gumbel Softmax
- Layer Flexible Adaptive Computational Time
- Unbiased Gradient Estimation with Balanced Assignments for Mixtures of Experts
- Sparsely ensembled convolutional neural network classifiers via reinforcement learning
- Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency
- Approximation Algorithms for Cascading Prediction Models
- Adaptive Attention Span in Transformers
- Sound Event Detection with Adaptive Frequency Selection
- Numerical Sequence Prediction using Bayesian Concept Learning
- Progress Extrapolating Algorithmic Learning to Arbitrary Sequence Lengths
- An investigation of model-free planning
- Training a Subsampling Mechanism in Expectation
- DA-LSTM: A Long Short-Term Memory with Depth Adaptive to Non-uniform Information Flow in Sequential Data
- Subword Language Model for Query Auto-Completion
- A Survey on Green Deep Learning
- Content-adaptive Representation Learning for Fast Image Super-resolution
- Consistent Accelerated Inference via Confident Adaptive Transformers