Feedforward Sequential Memory Networks: A New Structure to Learn Long-term Dependency
arXiv:1512.08301
Abstract
In this paper, we propose a novel neural network structure, namely \emph{feedforward sequential memory networks (FSMN)}, to model long-term dependency in time series without using recurrent feedback. The proposed FSMN is a standard fully-connected feedforward neural network equipped with some learnable memory blocks in its hidden layers. The memory blocks use a tapped-delay line structure to encode the long context information into a fixed-size representation as short-term memory mechanism. We have evaluated the proposed FSMNs in several standard benchmark tasks, including speech recognition and language modelling. Experimental results have shown FSMNs significantly outperform the conventional recurrent neural networks (RNN), including LSTMs, in modeling sequential signals like speech or language. Moreover, FSMNs can be learned much more reliably and faster than RNNs or LSTMs due to the inherent non-recurrent model structure.
11 pages, 5 figures
References in corpus (7)
- Deep Learning in Neural Networks: An Overview
- Neural Machine Translation by Jointly Learning to Align and Translate
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Recurrent Neural Network Regularization
- Feedforward Sequential Memory Neural Networks without Recurrent Feedback
- Hybrid Orthogonal Projection and Estimation (HOPE): A New Framework to Probe and Learn Neural Networks
Cited by in corpus (19)
- Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker Identification
- A novel pyramidal-FSMN architecture with lattice-free MMI for speech recognition
- Recognizing Multi-talker Speech with Permutation Invariant Training
- BiFSMN: Binary Neural Network for Keyword Spotting
- Universal ASR: Unifying Streaming and Non-Streaming ASR Using a Single Encoder-Decoder Model
- SAN-M: Memory Equipped Self-Attention for End-to-End Speech Recognition
- Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
- Advanced LSTM: A Study about Better Time Dependency Modeling in Emotion Recognition
- Empirical Evaluation of Speaker Adaptation on DNN based Acoustic Model
- Gated Recurrent Unit Based Acoustic Modeling with Future Context
- Model Interpolation with Trans-dimensional Random Field Language Models for Speech Recognition
- Recent Progresses in Deep Learning based Acoustic Models (Updated)
- End-to-End Streaming Keyword Spotting
- Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
- Deep Feed-forward Sequential Memory Networks for Speech Synthesis
- Improving Gated Recurrent Unit Based Acoustic Modeling with Batch Normalization and Enlarged Context
- Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
- Simplified Self-Attention for Transformer-based End-to-End Speech Recognition
- Residual Memory Networks: Feed-forward approach to learn long temporal dependencies