Highway Networks
arXiv:1505.00387
Abstract
There is plenty of theoretical and empirical evidence that depth of neural networks is a crucial ingredient for their success. However, network training becomes more difficult with increasing depth and training of very deep networks remains an open problem. In this extended abstract, we introduce a new architecture designed to ease gradient-based training of very deep networks. We refer to networks with this architecture as highway networks, since they allow unimpeded information flow across several layers on "information highways". The architecture is characterized by the use of gating units which learn to regulate the flow of information through a network. Highway networks with hundreds of layers can be trained directly using stochastic gradient descent and with a variety of activation functions, opening up the possibility of studying extremely deep and efficient architectures.
6 pages, 2 figures. Presented at ICML 2015 Deep Learning workshop. Full paper is at arXiv:1507.06228
Cited by in corpus (101)
- Deep Residual Learning for Image Recognition
- Wide Residual Networks
- Automatically designing CNN architectures using genetic algorithm for image classification
- Achieving Human Parity in Conversational Speech Recognition
- QANet: Combining Local Convolution with Global Self-Attention for Reading Comprehension
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- The Microsoft 2016 Conversational Speech Recognition System
- Efficient representation and approximation of model predictive control laws via deep learning
- Machine Learning for Wireless Communications in the Internet of Things: A Comprehensive Survey
- SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
- Neural GPUs Learn Algorithms
- Multi-view Graph Contrastive Representation Learning for Drug-Drug Interaction Prediction
- Tensor Methods in Computer Vision and Deep Learning
- FCN-Transformer Feature Fusion for Polyp Segmentation
- StegNet: Mega Image Steganography Capacity with Deep Convolutional Network
- Vision-based Real Estate Price Estimation
- Neural Networks for Entity Matching: A Survey
- Deep Residual Networks with Exponential Linear Unit
- Boosting the Speed of Entity Alignment 10*: Dual Attention Matching Network with Normalized Hard Sample Mining
- Discourse-Based Objectives for Fast Unsupervised Sentence Representation Learning
- Review: Deep Learning in Electron Microscopy
- The Shattered Gradients Problem: If resnets are the answer, then what is the question?
- Autoregressive Convolutional Neural Networks for Asynchronous Time Series
- Multi-modal fusion with gating using audio, lexical and disfluency features for Alzheimer's Dementia recognition from spontaneous speech
- Artificial Neural Networks for Photonic Applications: From Algorithms to Implementation
- Deep Polynomial Neural Networks
- Representation Learning for Natural Language Processing
- Scattering Networks for Hybrid Representation Learning
- MFRNet: A New CNN Architecture for Post-Processing and In-loop Filtering
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- AdS/CFT as a deep Boltzmann machine
- Learning Dynamic Belief Graphs to Generalize on Text-Based Games
- Unsupervised Deep Representation Learning and Few-Shot Classification of PolSAR Images
- Coordinated Sum-Rate Maximization in Multicell MU-MIMO with Deep Unrolling
- Dense and Diverse Capsule Networks: Making the Capsules Learn Better
- Evolutionary Preference Learning via Graph Nested GRU ODE for Session-based Recommendation
- Higher Order Recurrent Neural Networks
- Dual Co-Matching Network for Multi-choice Reading Comprehension
- Oriented Response Networks
- Deep Layer Aggregation
- Retrieve-and-Read: Multi-task Learning of Information Retrieval and Reading Comprehension
- Learning for Disparity Estimation through Feature Constancy
- Convolutional Spatial Attention Model for Reading Comprehension with Multiple-Choice Questions
- Is it enough to optimize CNN architectures on ImageNet?
- Don't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
- Neural Machine Reading Comprehension: Methods and Trends
- Data-Driven Sparse Structure Selection for Deep Neural Networks
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Multi-Cast Attention Networks for Retrieval-based Question Answering and Response Prediction
- Automatically Evolving CNN Architectures Based on Blocks
- Lightweight Residual Densely Connected Convolutional Neural Network
- QA4IE: A Question Answering based Framework for Information Extraction
- Reconstructing Patchy Reionization with Deep Learning
- Discriminative Sentence Modeling for Story Ending Prediction
- Training Deeper Neural Machine Translation Models with Transparent Attention
- Have convolutions already made recurrence obsolete for unconstrained handwritten text recognition ?
- Ruminating Reader: Reasoning with Gated Multi-Hop Attention
- Reconstructing Cosmic Polarization Rotation with ResUNet-CMB
- Interactive Language Learning by Question Answering
- Should You Go Deeper? Optimizing Convolutional Neural Network Architectures without Training by Receptive Field Analysis
- Combining Ensembles and Data Augmentation can Harm your Calibration
- Improving Interpretability of Word Embeddings by Generating Definition and Usage
- Express Wavenet -- a low parameter optical neural network with random shift wavelet pattern
- Learning to Multitask
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- The Shallow End: Empowering Shallower Deep-Convolutional Networks through Auxiliary Outputs
- Deep Cross Residual Learning for Multitask Visual Recognition
- Learning Strict Identity Mappings in Deep Residual Networks
- AOGNets: Compositional Grammatical Architectures for Deep Learning
- Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis
- Sparsely Aggregated Convolutional Networks
- Batch-normalized Recurrent Highway Networks
- Benchmarking Deep Sequential Models on Volatility Predictions for Financial Time Series
- SORT: Second-Order Response Transform for Visual Recognition
- Direct Output Connection for a High-Rank Language Model
- Supervised Deep Sparse Coding Networks
- SNDCNN: Self-normalizing deep CNNs with scaled exponential linear units for speech recognition
- Modeling Fine-Grained Entity Types with Box Embeddings
- Learning Filter Scale and Orientation In CNNs
- Balanced Binary Neural Networks with Gated Residual
- Automatic Inference of Cross-modal Connection Topologies for X-CNNs
- Learning Set-equivariant Functions with SWARM Mappings
- Predicting the Future with Transformational States
- High Order Recurrent Neural Networks for Acoustic Modelling
- Nonsymbolic Text Representation
- Interactive Machine Comprehension with Information Seeking Agents
- Toward Generalist Neural Motion Planners for Robotic Manipulators: Challenges and Opportunities
- Convolution in Convolution for Network in Network
- Tandem Blocks in Deep Convolutional Neural Networks
- Semi-tied Units for Efficient Gating in LSTM and Highway Networks
- Regularized Context Gates on Transformer for Machine Translation
- Object Recognition Based on Amounts of Unlabeled Data
- Simple2Complex: Global Optimization by Gradient Descent
- Rethinking Radiology: An Analysis of Different Approaches to BraTS
- Autoencoding Undirected Molecular Graphs With Neural Networks
- Efficient Rotation Invariance in Deep Neural Networks through Artificial Mental Rotation
- Sequentially Aggregated Convolutional Networks
- Adaptive Noise Injection: A Structure-Expanding Regularization for RNN
- Structure Learning of Deep Neural Networks with Q-Learning
- Using Domain Knowledge for Low Resource Named Entity Recognition
- Recurrent networks improve neural response prediction and provide insights into underlying cortical circuits