A Clockwork RNN
arXiv:1402.3511
Abstract
Sequence prediction and classification are ubiquitous and challenging problems in machine learning that can require identifying complex dependencies between temporally distant inputs. Recurrent Neural Networks (RNNs) have the ability, in theory, to cope with these temporal dependencies by virtue of the short-term memory implemented by their recurrent (feedback) connections. However, in practice they are difficult to train successfully when the long-term memory is required. This paper introduces a simple, yet powerful modification to the standard RNN architecture, the Clockwork RNN (CW-RNN), in which the hidden layer is partitioned into separate modules, each processing inputs at its own temporal granularity, making computations only at its prescribed clock rate. Rather than making the standard RNN models more complex, CW-RNN reduces the number of RNN parameters, improves the performance significantly in the tasks tested, and speeds up the network evaluation. The network is demonstrated in preliminary experiments involving two tasks: audio signal generation and TIMIT spoken word classification, where it outperforms both RNN and LSTM networks.
Cited by in corpus (55)
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Recurrent Neural Networks for Time Series Forecasting: Current Status and Future Directions
- Human-level performance in first-person multiplayer games with population-based deep reinforcement learning
- End-To-End Memory Networks
- Listen, Attend and Spell
- Learning to Execute
- Neural Granger Causality
- Long Short-Term Memory-Networks for Machine Reading
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Understanding LSTM -- a tutorial into Long Short-Term Memory Recurrent Neural Networks
- Learning Contextual Dependencies with Convolutional Hierarchical Recurrent Neural Networks
- Architectural Complexity Measures of Recurrent Neural Networks
- Learning Longer-term Dependencies in RNNs with Auxiliary Losses
- Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks
- Recurrent Independent Mechanisms
- Improving performance of recurrent neural network with relu nonlinearity
- Variational Temporal Deep Generative Model for Radar HRRP Target Recognition
- Real-Time Well Log Prediction From Drilling Data Using Deep Learning
- ST-UNet: A Spatio-Temporal U-Network for Graph-structured Time Series Modeling
- From Fourier to Koopman: Spectral Methods for Long-term Time Series Prediction
- Time-series modeling with undecimated fully convolutional neural networks
- Higher Order Recurrent Neural Networks
- Predictive-Corrective Networks for Action Detection
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- On Learning to Think: Algorithmic Information Theory for Novel Combinations of Reinforcement Learning Controllers and Recurrent Neural World Models
- Grouped Convolutional Neural Networks for Multivariate Time Series
- Efficient Orthogonal Parametrisation of Recurrent Neural Networks Using Householder Reflections
- Improving the Gating Mechanism of Recurrent Neural Networks
- Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling
- Discovering Signals from Web Sources to Predict Cyber Attacks
- Focused Hierarchical RNNs for Conditional Sequence Processing
- Neural Language Modeling by Jointly Learning Syntax and Lexicon
- Analyzing and Exploiting NARX Recurrent Neural Networks for Long-Term Dependencies
- Bidirectional Recurrent Neural Networks as Generative Models - Reconstructing Gaps in Time Series
- Sliced Recurrent Neural Networks
- Multi-variable LSTM neural network for autoregressive exogenous model
- A Hierarchical Approach for Generating Descriptive Image Paragraphs
- Recurrent Neural Processes
- Temporal Difference Variational Auto-Encoder
- From Artificial Neural Networks to Deep Learning for Music Generation -- History, Concepts and Trends
- A Hierarchical Recurrent Neural Network for Symbolic Melody Generation
- Multilevel Wavelet Decomposition Network for Interpretable Time Series Analysis
- Cell-aware Stacked LSTMs for Modeling Sentences
- Neural Pharmacodynamic State Space Modeling
- Massively Parallel Video Networks
- Low-pass Recurrent Neural Networks - A memory architecture for longer-term correlation discovery
- Deep Learning: Our Miraculous Year 1990-1991
- Learning Simpler Language Models with the Differential State Framework
- Better Long-Range Dependency By Bootstrapping A Mutual Information Regularizer
- Deep Recurrent Encoder: A scalable end-to-end network to model brain signals
- Sequence Prediction using Spectral RNNs
- Linked Recurrent Neural Networks
- A Hierarchical Transformer for Unsupervised Parsing
- Variational Predictive Routing with Nested Subjective Timescales
- Compositional Sentence Representation from Character within Large Context Text