Hierarchical Multiscale Recurrent Neural Networks
arXiv:1609.01704
Abstract
Learning both hierarchical and temporal representation has been among the long-standing challenges of recurrent neural networks. Multiscale recurrent neural networks have been considered as a promising approach to resolve this issue, yet there has been a lack of empirical evidence showing that this type of models can actually capture the temporal dependencies by discovering the latent hierarchical structure of the sequence. In this paper, we propose a novel multiscale approach, called the hierarchical multiscale recurrent neural networks, which can capture the latent hierarchical structure in the sequence by encoding the temporal dependencies with different timescales using a novel update mechanism. We show some evidence that our proposed multiscale architecture can discover underlying hierarchical structure in the sequences without using explicit boundary information. We evaluate our proposed model on character-level language modelling and handwriting sequence modelling.
References in corpus (4)
Cited by in corpus (39)
- Neural Machine Translation in Linear Time
- Learning Multimodal Graph-to-Graph Translation for Molecular Optimization
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- Lifelong Sequential Modeling with Personalized Memorization for User Response Prediction
- Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
- Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
- Dynamic Evaluation of Neural Sequence Models
- Compressive Transformers for Long-Range Sequence Modelling
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Fast-Slow Recurrent Neural Networks
- SyntaxNet Models for the CoNLL 2017 Shared Task
- 3G structure for image caption generation
- Deep-ESN: A Multiple Projection-encoding Hierarchical Reservoir Computing Framework
- Deep Learning for Multi-Scale Changepoint Detection in Multivariate Time Series
- Event Identification as a Decision Process with Non-linear Representation of Text
- Unsupervised Video Decomposition using Spatio-temporal Iterative Inference
- Learning to Attend, Copy, and Generate for Session-Based Query Suggestion
- Surprisal-Driven Zoneout
- Multiscale sequence modeling with a learned dictionary
- MEMO: A Deep Network for Flexible Combination of Episodic Memories
- Hide-and-Seek: A Template for Explainable AI
- Hierarchical Multi-scale Attention Networks for Action Recognition
- Multi-Scale Self-Attention for Text Classification
- Stochastic Sequential Neural Networks with Structured Inference
- Multi-scale Transformer Language Models
- Encoding-based Memory Modules for Recurrent Neural Networks
- Multi-Zone Unit for Recurrent Neural Networks
- Learning to Adaptively Scale Recurrent Neural Networks
- Rotational Unit of Memory
- Shortcut Sequence Tagging
- Perception-Prediction-Reaction Agents for Deep Reinforcement Learning
- Options Discovery with Budgeted Reinforcement Learning
- Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video Data
- Decoupling Hierarchical Recurrent Neural Networks With Locally Computable Losses
- Solving Large-Scale 0-1 Knapsack Problems and its Application to Point Cloud Resampling
- ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network
- Analysis of memory in LSTM-RNNs for source separation
- DA-LSTM: A Long Short-Term Memory with Depth Adaptive to Non-uniform Information Flow in Sequential Data
- A memory enhanced LSTM for modeling complex temporal dependencies