Visualizing and Understanding Recurrent Networks
arXiv:1506.02078
Abstract
Recurrent Neural Networks (RNNs), and specifically a variant with Long Short-Term Memory (LSTM), are enjoying renewed interest as a result of successful applications in a wide range of machine learning problems that involve sequential data. However, while LSTMs provide exceptional results in practice, the source of their performance and their limitations remain rather poorly understood. Using character-level language models as an interpretable testbed, we aim to bridge this gap by providing an analysis of their representations, predictions and error types. In particular, our experiments reveal the existence of interpretable cells that keep track of long-range dependencies such as line lengths, quotes and brackets. Moreover, our comparative analysis with finite horizon n-gram models traces the source of the LSTM improvements to long-range structural dependencies. Finally, we provide analysis of the remaining errors and suggests areas for further study.
changing style, adding references, minor changes to text
References in corpus (5)
Cited by in corpus (262)
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- Interpretable machine learning: definitions, methods, and applications
- DeepSleepNet: a Model for Automatic Sleep Stage Scoring based on Raw Single-Channel EEG
- A trans-disciplinary review of deep learning research for water resources scientists
- Deep Reinforcement Learning: An Overview
- Deep Neural Networks for Bot Detection
- Understanding Neural Networks through Representation Erasure
- Predicting Process Behaviour using Deep Learning
- Understanding the Role of Individual Units in a Deep Neural Network
- NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis
- Learning to Generate Reviews and Discovering Sentiment
- Deep Learning for Sensor-based Human Activity Recognition: Overview, Challenges and Opportunities
- RetainVis: Visual Analytics with Interpretable and Interactive Recurrent Neural Networks on Electronic Medical Records
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- What do Neural Machine Translation Models Learn about Morphology?
- Doctor AI: Predicting Clinical Events via Recurrent Neural Networks
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- RETAIN: An Interpretable Predictive Model for Healthcare using Reverse Time Attention Mechanism
- Predicting Domain Generation Algorithms with Long Short-Term Memory Networks
- On Extended Long Short-term Memory and Dependent Bidirectional Recurrent Neural Network
- GAN Dissection: Visualizing and Understanding Generative Adversarial Networks
- Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
- Visualizing and Understanding Neural Models in NLP
- The Neural Hawkes Process: A Neurally Self-Modulating Multivariate Point Process
- Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks
- Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model
- NeuralHydrology -- Interpreting LSTMs in Hydrology
- TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
- RSDNet: Learning to Predict Remaining Surgery Duration from Laparoscopic Videos Without Manual Annotations
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- On the Explainability of Natural Language Processing Deep Models
- Language to Logical Form with Neural Attention
- Recent Advances in Deep Learning: An Overview
- Machine Learning Testing: Survey, Landscapes and Horizons
- Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
- Recurrent Dropout without Memory Loss
- Natural Language Statistical Features of LSTM-generated Texts
- Deep Active Inference
- Loss-aware Weight Quantization of Deep Networks
- Explainable Artificial Intelligence: a Systematic Review
- Peephole: Predicting Network Performance Before Training
- DenseCap: Fully Convolutional Localization Networks for Dense Captioning
- Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks
- Tree-to-tree Neural Networks for Program Translation
- Noisy Activation Functions
- Recurrent Neural Networks for Time Series Forecasting
- Gated Siamese Convolutional Neural Network Architecture for Human Re-Identification
- Context-aware Natural Language Generation with Recurrent Neural Networks
- High-Accuracy Low-Precision Training
- Improving Deep Learning for HAR with shallow LSTMs
- Recursive Recurrent Nets with Attention Modeling for OCR in the Wild
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
- A Siamese Long Short-Term Memory Architecture for Human Re-Identification
- Cognitive Psychology for Deep Neural Networks: A Shape Bias Case Study
- Simple and Accurate Dependency Parsing Using Bidirectional LSTM Feature Representations
- Insights on representational similarity in neural networks with canonical correlation
- Increasing the Interpretability of Recurrent Neural Networks Using Hidden Markov Models
- Tree-structured composition in neural networks without tree-structured architectures
- The emergence of number and syntax units in LSTM language models
- On the Binding Problem in Artificial Neural Networks
- deepMiRGene: Deep Neural Network based Precursor microRNA Prediction
- On Interpretability of Artificial Neural Networks: A Survey
- Machine learning astrophysics from 21 cm lightcones: impact of network architectures and signal contamination
- Explain and Predict, and then Predict Again
- Interpreting a Recurrent Neural Network's Predictions of ICU Mortality Risk
- Towards Binary-Valued Gates for Robust LSTM Training
- Dual-Branched Spatio-temporal Fusion Network for Multi-horizon Tropical Cyclone Track Forecast
- Predictive-Corrective Networks for Action Detection
- DRLViz: Understanding Decisions and Memory in Deep Reinforcement Learning
- Interpretable Recurrent Neural Networks Using Sequential Sparse Recovery
- Linguistic Profiling of a Neural Language Model
- Dank Learning: Generating Memes Using Deep Neural Networks
- Implicit Language Model in LSTM for OCR
- Short-term traffic prediction using physics-aware neural networks
- Learning to Read Chest X-Rays: Recurrent Neural Cascade Model for Automated Image Annotation
- A Taxonomy for Neural Memory Networks
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution
- Detecting Statistical Interactions from Neural Network Weights
- Multi-modal gated recurrent units for image description
- AI Safety for Everyone
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Predicting purchasing intent: Automatic Feature Learning using Recurrent Neural Networks
- Identifying and Controlling Important Neurons in Neural Machine Translation
- Visualizing and Understanding Curriculum Learning for Long Short-Term Memory Networks
- Towards falsifiable interpretability research
- GenNI: Human-AI Collaboration for Data-Backed Text Generation
- State Space LSTM Models with Particle MCMC Inference
- Applying Deep Learning to Basketball Trajectories
- Techniques for Interpretable Machine Learning
- Backward and Forward Language Modeling for Constrained Sentence Generation
- Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
- Reverse engineering recurrent networks for sentiment classification reveals line attractor dynamics
- Explaining Time Series Predictions with Dynamic Masks
- Understanding Hidden Memories of Recurrent Neural Networks
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Self-Explaining Structures Improve NLP Models
- Deep Motif Dashboard: Visualizing and Understanding Genomic Sequences Using Deep Neural Networks
- Towards Frequency-Based Explanation for Robust CNN
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- Learning the Enigma with Recurrent Neural Networks
- Task-driven Visual Saliency and Attention-based Visual Question Answering
- Recurrent Memory Networks for Language Modeling
- Large-Scale User Modeling with Recurrent Neural Networks for Music Discovery on Multiple Time Scales
- Self-Attentive Hawkes Processes
- Deep Learning in Bioinformatics
- Selectivity considered harmful: evaluating the causal impact of class selectivity in DNNs
- How do Mixture Density RNNs Predict the Future?
- Seeing the Wind: Visual Wind Speed Prediction with a Coupled Convolutional and Recurrent Neural Network
- Semantics and explanation: why counterfactual explanations produce adversarial examples in deep neural networks
- MinimalRNN: Toward More Interpretable and Trainable Recurrent Neural Networks
- Improving Interpretability of Deep Neural Networks with Semantic Information
- Interpretable Deep Learning under Fire
- xGEMs: Generating Examplars to Explain Black-Box Models
- Deep reinforcement learning for time series: playing idealized trading games
- Open Sesame: Getting Inside BERT's Linguistic Knowledge
- Human Motion Prediction via Spatio-Temporal Inpainting
- Gate Activation Signal Analysis for Gated Recurrent Neural Networks and Its Correlation with Phoneme Boundaries
- Representation of linguistic form and function in recurrent neural networks
- On Learning and Learned Data Representation by Capsule Networks
- LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Strongly-Typed Recurrent Neural Networks
- AllenNLP Interpret: A Framework for Explaining Predictions of NLP Models
- Mediators in Determining what Processing BERT Performs First
- Techniques for visualizing LSTMs applied to electrocardiograms
- Toward a Theory of Causation for Interpreting Neural Code Models
- Bridging LSTM Architecture and the Neural Dynamics during Reading
- End-to-End Prediction of Buffer Overruns from Raw Source Code via Neural Memory Networks
- How recurrent networks implement contextual processing in sentiment analysis
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- Demystifying Deep Learning in Predictive Spatio-Temporal Analytics: An Information-Theoretic Framework
- On the Units of GANs (Extended Abstract)
- Compressing Neural Language Models by Sparse Word Representations
- What Do Recurrent Neural Network Grammars Learn About Syntax?
- Why Do Neural Dialog Systems Generate Short and Meaningless Replies? A Comparison between Dialog and Translation
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Learning Visual Storylines with Skipping Recurrent Neural Networks
- Neural-Symbolic Reasoning over Knowledge Graph for Multi-stage Explainable Recommendation
- Semi Supervised Preposition-Sense Disambiguation using Multilingual Data
- Visualizing and Understanding Sum-Product Networks
- Stability of Internal States in Recurrent Neural Networks Trained on Regular Languages
- Mutual Information Scaling and Expressive Power of Sequence Models
- Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
- On the Interpretability of Deep Learning Based Models for Knowledge Tracing
- Can LSTM Learn to Capture Agreement? The Case of Basque
- Conditioning Deep Generative Raw Audio Models for Structured Automatic Music
- Machine Learning Techniques for Software Quality Assurance: A Survey
- Improving Explainable Recommendations with Synthetic Reviews
- Learning Operations on a Stack with Neural Turing Machines
- Syntactically Informed Text Compression with Recurrent Neural Networks
- Dance Dance Convolution
- ProtoryNet - Interpretable Text Classification Via Prototype Trajectories
- Cached Long Short-Term Memory Neural Networks for Document-Level Sentiment Classification
- Deep Learning in Information Security
- Dense Image Representation with Spatial Pyramid VLAD Coding of CNN for Locally Robust Captioning
- DeepSeer: Interactive RNN Explanation and Debugging via State Abstraction
- X-ToM: Explaining with Theory-of-Mind for Gaining Justified Human Trust
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Learning higher-order sequential structure with cloned HMMs
- Top-down Visual Saliency Guided by Captions
- A journey in ESN and LSTM visualisations on a language task
- Formal Language Theory Meets Modern NLP
- Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark
- Learning with Interpretable Structure from Gated RNN
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Attentive Action and Context Factorization
- LS-Tree: Model Interpretation When the Data Are Linguistic
- Layerwise Knowledge Extraction from Deep Convolutional Networks
- M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis
- exploRNN: Understanding Recurrent Neural Networks through Visual Exploration
- Recurrent Memory Array Structures
- Energy-Based Models for Code Generation under Compilability Constraints
- A Survey on Understanding, Visualizations, and Explanation of Deep Neural Networks
- Exemplary Natural Images Explain CNN Activations Better than State-of-the-Art Feature Visualization
- Visualizing RNN States with Predictive Semantic Encodings
- Discovery of Natural Language Concepts in Individual Units of CNNs
- The geometry of integration in text classification RNNs
- Visualizing textual models with in-text and word-as-pixel highlighting
- Towards More Fine-grained and Reliable NLP Performance Prediction
- On the relationship between class selectivity, dimensionality, and robustness
- Distance and Equivalence between Finite State Machines and Recurrent Neural Networks: Computational results
- A Maximum Matching Algorithm for Basis Selection in Spectral Learning
- Interpretable Text Classification Using CNN and Max-pooling
- DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
- Security Vulnerability Detection Using Deep Learning Natural Language Processing
- GRACE: Generating Concise and Informative Contrastive Sample to Explain Neural Network Model's Prediction
- Does it care what you asked? Understanding Importance of Verbs in Deep Learning QA System
- Learning Recurrent Binary/Ternary Weights
- EMAP: Explanation by Minimal Adversarial Perturbation
- Analysis Methods in Neural Language Processing: A Survey
- Low-Rank RNN Adaptation for Context-Aware Language Modeling
- Improving Context Aware Language Models
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
- Distinguishing rule- and exemplar-based generalization in learning systems
- DeepEverest: Accelerating Declarative Top-K Queries for Deep Neural Network Interpretation
- Investigating how well contextual features are captured by bi-directional recurrent neural network models
- Sampling for Deep Learning Model Diagnosis (Technical Report)
- Multi-Task Spatiotemporal Neural Networks for Structured Surface Reconstruction
- FIND: Human-in-the-Loop Debugging Deep Text Classifiers
- Neural Machine Translation: A Review and Survey
- Generating Philosophical Statements using Interpolated Markov Models and Dynamic Templates
- Understanding and Controlling Memory in Recurrent Neural Networks
- Analyzing and Interpreting Convolutional Neural Networks in NLP
- Applications of Probabilistic Programming (Master's thesis, 2015)
- Visualizing MuZero Models
- Learning Music Helps You Read: Using Transfer to Study Linguistic Structure in Language Models
- Do RNN States Encode Abstract Phonological Processes?
- Persistence pays off: Paying Attention to What the LSTM Gating Mechanism Persists
- Exploring the Naturalness of Buggy Code with Recurrent Neural Networks
- A Simple Recurrent Unit with Reduced Tensor Product Representations
- Joint learning of interpretation and distillation
- Understanding Recurrent Neural State Using Memory Signatures
- Non-Projective Dependency Parsing via Latent Heads Representation (LHR)
- Simplifying the explanation of deep neural networks with sufficient and necessary feature-sets: case of text classification
- Assessing the Memory Ability of Recurrent Neural Networks
- Interpreting Deep Learning Model Using Rule-based Method
- Formal models of Structure Building in Music, Language and Animal Songs
- LAMVI-2: A Visual Tool for Comparing and Tuning Word Embedding Models
- Deep Learning Scooping Motion using Bilateral Teleoperations
- Forced to Learn: Discovering Disentangled Representations Without Exhaustive Labels
- Probabilistic Modeling for Novelty Detection with Applications to Fraud Identification
- On Attribution of Recurrent Neural Network Predictions via Additive Decomposition
- Generalized Constraints as A New Mathematical Problem in Artificial Intelligence: A Review and Perspective
- A Comparative Analysis of Knowledge-Intensive and Data-Intensive Semantic Parsers
- Explain by Evidence: An Explainable Memory-based Neural Network for Question Answering
- Inverting and Understanding Object Detectors
- TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes Applications
- Visualizing Deep Learning-based Radio Modulation Classifier
- When and where do feed-forward neural networks learn localist representations?
- Every Filter Extracts A Specific Texture In Convolutional Neural Networks
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Interpreting A Pre-trained Model Is A Key For Model Architecture Optimization: A Case Study On Wav2Vec 2.0
- Influence Paths for Characterizing Subject-Verb Number Agreement in LSTM Language Models
- Sparsity Emerges Naturally in Neural Language Models
- Variable-sized input, character-level recurrent neural networks in lead generation: predicting close rates from raw user inputs
- Char-RNN and Active Learning for Hashtag Segmentation
- A2: Extracting Cyclic Switchings from DOB-nets for Rejecting Excessive Disturbances
- Do WaveNets Dream of Acoustic Waves?
- Efficient Modelling Across Time of Human Actions and Interactions
- Understanding Memory Modules on Learning Simple Algorithms
- Visual Summary of Value-level Feature Attribution in Prediction Classes with Recurrent Neural Networks
- Applying Incremental Deep Neural Networks-based Posture Recognition Model for Injury Risk Assessment in Construction
- Improving Distributed Representations of Tweets - Present and Future
- Neural Supervised Domain Adaptation by Augmenting Pre-trained Models with Random Units
- Improving Moderation of Online Discussions via Interpretable Neural Models
- An Embedded Deep Learning based Word Prediction
- How LSTM Encodes Syntax: Exploring Context Vectors and Semi-Quantization on Natural Text
- Recurrent Neural Networks based Obesity Status Prediction Using Activity Data
- Data-Based Models for Hurricane Evolution Prediction: A Deep Learning Approach
- Explainable Adversarial Attacks in Deep Neural Networks Using Activation Profiles
- Are there any 'object detectors' in the hidden layers of CNNs trained to identify objects or scenes?
- Technical notes: Syntax-aware Representation Learning With Pointer Networks
- Phoneme Level Language Models for Sequence Based Low Resource ASR
- Fine-grained Interpretation and Causation Analysis in Deep NLP Models
- Difference-in-Differences: Bridging Normalization and Disentanglement in PG-GAN
- Progress Estimation and Phase Detection for Sequential Processes
- i-Algebra: Towards Interactive Interpretability of Deep Neural Networks
- Sensei: Self-Supervised Sensor Name Segmentation
- Rethinking the Form of Latent States in Image Captioning
- Linking average- and worst-case perturbation robustness via class selectivity and dimensionality
- Can a Compact Neuronal Circuit Policy be Re-purposed to Learn Simple Robotic Control?