Training Very Deep Networks
arXiv:1507.06228
Abstract
Theoretical and empirical evidence indicates that the depth of neural networks is crucial for their success. However, training becomes more difficult as depth increases, and training of very deep networks remains an open problem. Here we introduce a new architecture designed to overcome this. Our so-called highway networks allow unimpeded information flow across many layers on information highways. They are inspired by Long Short-Term Memory recurrent networks and use adaptive gating units to regulate the information flow. Even with hundreds of layers, highway networks can be trained directly through simple gradient descent. This enables the study of extremely deep and efficient architectures.
11 pages. Extends arXiv:1505.00387. Project webpage is at http://people.idsia.ch/~rupesh/very_deep_learning/. in Advances in Neural Information Processing Systems 2015
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Grid Long Short-Term Memory
- Spatially-sparse convolutional neural networks
- Random Walk Initialization for Training Very Deep Feedforward Networks
- On the Expressive Efficiency of Sum Product Networks
- Binding via Reconstruction Clustering
Cited by in corpus (265)
- Deep Residual Learning for Image Recognition
- Attention Mechanisms in Computer Vision: A Survey
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- DGM: A deep learning algorithm for solving partial differential equations
- Ensemble deep learning: A review
- Learning to Plan Chemical Syntheses
- Conditional Image Generation with PixelCNN Decoders
- Low-Dose CT with a Residual Encoder-Decoder Convolutional Neural Network (RED-CNN)
- Densely Connected Convolutional Networks
- Accurate De Novo Prediction of Protein Contact Map by Ultra-Deep Learning Model
- Character-Aware Neural Language Models
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- Automatically designing CNN architectures using genetic algorithm for image classification
- Exploring the Limits of Language Modeling
- Generalizing from a Few Examples: A Survey on Few-Shot Learning
- Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification
- Learning Representations by Maximizing Mutual Information Across Views
- Resnet in Resnet: Generalizing Residual Architectures
- Automatic lung segmentation in routine imaging is primarily a data diversity problem, not a methodology problem
- Group Equivariant Convolutional Networks
- Fully Dense UNet for 2D Sparse Photoacoustic Tomography Artifact Removal
- Large-Scale Evolution of Image Classifiers
- Training Deep Nets with Sublinear Memory Cost
- Dynamic Coattention Networks For Question Answering
- AMP-Inspired Deep Networks for Sparse Linear Inverse Problems
- Dynamic Filter Networks
- Jointly Multiple Events Extraction via Attention-based Graph Information Aggregation
- Deep Learning Based Text Classification: A Comprehensive Review
- DC-SPP-YOLO: Dense Connection and Spatial Pyramid Pooling Based YOLO for Object Detection
- Adaptive Computation Time for Recurrent Neural Networks
- Recent Advances in Convolutional Neural Networks
- Residual Attention Network for Image Classification
- Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention
- Deep 1D-Convnet for accurate Parkinson disease detection and severity prediction from gait
- An overview and comparative analysis of Recurrent Neural Networks for Short Term Load Forecasting
- Deep Networks with Stochastic Depth
- Single Model Deep Learning on Imbalanced Small Datasets for Skin Lesion Classification
- Adaptive Aggregation Networks for Class-Incremental Learning
- Stacked What-Where Auto-encoders
- Spatial-Spectral Feature Extraction via Deep ConvLSTM Neural Networks for Hyperspectral Image Classification
- Identity Mappings in Deep Residual Networks
- All you need is a good init
- Learning Depth from Single Images with Deep Neural Network Embedding Focal Length
- A Comprehensive Study of Deep Bidirectional LSTM RNNs for Acoustic Modeling in Speech Recognition
- Deep Complex Networks
- Understanding and Improving Convolutional Neural Networks via Concatenated Rectified Linear Units
- Discrimination-aware Channel Pruning for Deep Neural Networks
- Bridging Category-level and Instance-level Semantic Image Segmentation
- Discrimination-aware Network Pruning for Deep Model Compression
- Vision-based Real Estate Price Estimation
- Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation
- The Modern Mathematics of Deep Learning
- Swapout: Learning an ensemble of deep architectures
- Deep Residual Networks with Exponential Linear Unit
- Review: Deep Learning in Electron Microscopy
- Do Deep Convolutional Nets Really Need to be Deep and Convolutional?
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Deep Enhanced Representation for Implicit Discourse Relation Recognition
- High-performance Semantic Segmentation Using Very Deep Fully Convolutional Networks
- Overcoming the vanishing gradient problem in plain recurrent networks
- A Survey of Complex-Valued Neural Networks
- MFRNet: A New CNN Architecture for Post-Processing and In-loop Filtering
- Analyzing Human-Human Interactions: A Survey
- Very Deep Transformers for Neural Machine Translation
- The unreasonable effectiveness of the forget gate
- Building medical image classifiers with very limited data using segmentation networks
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- Deconstructing the Ladder Network Architecture
- RFBNet: Deep Multimodal Networks with Residual Fusion Blocks for RGB-D Semantic Segmentation
- Deep Motif: Visualizing Genomic Sequence Classifications
- Convolutional Neural Networks with Gated Recurrent Connections
- Transformation Networks for Target-Oriented Sentiment Classification
- 6GAN: IPv6 Multi-Pattern Target Generation via Generative Adversarial Nets with Reinforcement Learning
- Deciding How to Decide: Dynamic Routing in Artificial Neural Networks
- On Warm-Starting Neural Network Training
- Deep Predictive Coding Network for Object Recognition
- Inspect, Understand, Overcome: A Survey of Practical Methods for AI Safety
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- Multi-layer Representation Fusion for Neural Machine Translation
- Single Image Super-Resolution via Cascaded Multi-Scale Cross Network
- IRNet: A General Purpose Deep Residual Regression Framework for Materials Discovery
- Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
- Noisy Differentiable Architecture Search
- SketchyGAN: Towards Diverse and Realistic Sketch to Image Synthesis
- PolyNet: A Pursuit of Structural Diversity in Very Deep Networks
- Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Quantum-inspired Complex Convolutional Neural Networks
- GFF: Gated Fully Fusion for Semantic Segmentation
- Automatic Speech Recognition with Very Large Conversational Finnish and Estonian Vocabularies
- Automated curricula through setter-solver interactions
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary Search
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? -- A Neural Tangent Kernel Perspective
- Feature Fusion for Online Mutual Knowledge Distillation
- Label Embedding Network: Learning Label Representation for Soft Training of Deep Networks
- Deep Pyramidal Residual Networks
- CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters
- End-to-End Learning for Structured Prediction Energy Networks
- Explicit Contextual Semantics for Text Comprehension
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- Regularizing Deep Networks with Semantic Data Augmentation
- Guardians of the Deep Fog: Failure-Resilient DNN Inference from Edge to Cloud
- SAUNet: Shape Attentive U-Net for Interpretable Medical Image Segmentation
- Lightweight Residual Densely Connected Convolutional Neural Network
- End-to-end Deep Learning from Raw Sensor Data: Atrial Fibrillation Detection using Wearables
- Advanced Capsule Networks via Context Awareness
- ANTNets: Mobile Convolutional Neural Networks for Resource Efficient Image Classification
- Low-resource Deep Entity Resolution with Transfer and Active Learning
- Character-Word LSTM Language Models
- Single-epoch supernova classification with deep convolutional neural networks
- Selfish Sparse RNN Training
- TTS-Portuguese Corpus: a corpus for speech synthesis in Brazilian Portuguese
- Fitting New Speakers Based on a Short Untranscribed Sample
- Not to Cry Wolf: Distantly Supervised Multitask Learning in Critical Care
- Layer Pruning via Fusible Residual Convolutional Block for Deep Neural Networks
- Deep Watershed Detector for Music Object Recognition
- Deep Neural Machine Translation with Linear Associative Unit
- VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
- Attention-Based Convolutional Neural Network for Machine Comprehension
- Dense Dilated UNet: Deep Learning for 3D Photoacoustic Tomography Image Reconstruction
- Parallelizing Legendre Memory Unit Training
- Genetic Architect: Discovering Genomic Structure with Learned Neural Architectures
- Learning Short-Cut Connections for Object Counting
- Improving Deep Neural Network with Multiple Parametric Exponential Linear Units
- On Deep Instrumental Variables Estimate
- DeNet: Scalable Real-time Object Detection with Directed Sparse Sampling
- TLDR: Token Loss Dynamic Reweighting for Reducing Repetitive Utterance Generation
- The effect of Target Normalization and Momentum on Dying ReLU
- Exploring Self-Supervised Regularization for Supervised and Semi-Supervised Learning
- CU-Net: Coupled U-Nets
- Learning in the Machine: Random Backpropagation and the Deep Learning Channel
- Convolutional Networks with Dense Connectivity
- A Unified Speaker Adaptation Method for Speech Synthesis using Transcribed and Untranscribed Speech with Backpropagation
- Shape-from-Mask: A Deep Learning Based Human Body Shape Reconstruction from Binary Mask Images
- MotionRNN: A Flexible Model for Video Prediction with Spacetime-Varying Motions
- Text Summarization with Latent Queries
- Fitted Learning: Models with Awareness of their Limits
- The Shallow End: Empowering Shallower Deep-Convolutional Networks through Auxiliary Outputs
- Blending LSTMs into CNNs
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Thinking Globally, Acting Locally: Distantly Supervised Global-to-Local Knowledge Selection for Background Based Conversation
- An evaluation of randomized machine learning methods for redundant data: Predicting short and medium-term suicide risk from administrative records and risk assessments
- Word Shape Matters: Robust Machine Translation with Visual Embedding
- Syntactic Scaffolds for Semantic Structures
- A^2-Net: Molecular Structure Estimation from Cryo-EM Density Volumes
- Neural-networks-based Photon-Counting Data Correction: Pulse Pileup Effect
- Parallel Extraction of Long-term Trends and Short-term Fluctuation Framework for Multivariate Time Series Forecasting
- Attentional Feature Fusion
- HarDNet: A Low Memory Traffic Network
- Faster Convergence in Deep-Predictive-Coding Networks to Learn Deeper Representations
- Detecting The Objects on The Road Using Modular Lightweight Network
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Why Adversarial Reprogramming Works, When It Fails, and How to Tell the Difference
- Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech
- Learning From Less Data: Diversified Subset Selection and Active Learning in Image Classification Tasks
- Densely Connected Bidirectional LSTM with Applications to Sentence Classification
- On the Depth of Deep Neural Networks: A Theoretical View
- Cascaded Subpatch Networks for Effective CNNs
- Attentional Multilabel Learning over Graphs: A Message Passing Approach
- NatCSNN: A Convolutional Spiking Neural Network for recognition of objects extracted from natural images
- Demystifying ResNet
- Improving Sequential Latent Variable Models with Autoregressive Flows
- DeepFreak: Learning Crystallography Diffraction Patterns with Automated Machine Learning
- On the Pitfalls of Learning with Limited Data: A Facial Expression Recognition Case Study
- Multiresolution Transformer Networks: Recurrence is Not Essential for Modeling Hierarchical Structure
- Deep Kernel Survival Analysis and Subject-Specific Survival Time Prediction Intervals
- Factorized Bilinear Models for Image Recognition
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- An Empirical Study of Language CNN for Image Captioning
- Iterative Amortized Policy Optimization
- SORT: Second-Order Response Transform for Visual Recognition
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- Open DNN Box by Power Side-Channel Attack
- Mollifying Networks
- Training CNNs with Selective Allocation of Channels
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- Bootstrapping NLU Models with Multi-task Learning
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- Syntax-aware Neural Semantic Role Labeling
- Let's be Humorous: Knowledge Enhanced Humor Generation
- Learning Channel Inter-dependencies at Multiple Scales on Dense Networks for Face Recognition
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- Attentive Convolution: Equipping CNNs with RNN-style Attention Mechanisms
- Graph Highway Networks
- Neural Style Representations and the Large-Scale Classification of Artistic Style
- Exploiting Contextual Information with Deep Neural Networks
- Unsupervised Machine Learning Discovery of Chemical and Physical Transformation Pathways from Imaging Data
- Residual Encoder-Decoder Network for Deep Subspace Clustering
- Refined Gate: A Simple and Effective Gating Mechanism for Recurrent Units
- Choice by Elimination via Deep Neural Networks
- HGC: Hierarchical Group Convolution for Highly Efficient Neural Network
- Learning Latent Causal Structures with a Redundant Input Neural Network
- On a Sparse Shortcut Topology of Artificial Neural Networks
- Effective Subword Segmentation for Text Comprehension
- Tensor Switching Networks
- Deep Context-Aware Kernel Networks
- Exploring Weight Symmetry in Deep Neural Networks
- Deep Decomposition Learning for Inverse Imaging Problems
- Initializing ReLU networks in an expressive subspace of weights
- Empirical Frequentist Coverage of Deep Learning Uncertainty Quantification Procedures
- Smaller Models, Better Generalization
- A Response Retrieval Approach for Dialogue Using a Multi-Attentive Transformer
- Alignment Attention by Matching Key and Query Distributions
- The Deep Parametric PDE Method: Application to Option Pricing
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- Geometrically Principled Connections in Graph Neural Networks
- RecNets: Channel-wise Recurrent Convolutional Neural Networks
- Accent Estimation of Japanese Words from Their Surfaces and Romanizations for Building Large Vocabulary Accent Dictionaries
- Attention-Based Clustering: Learning a Kernel from Context
- SuperNet -- An efficient method of neural networks ensembling
- MaskParse@Deskin at SemEval-2019 Task 1: Cross-lingual UCCA Semantic Parsing using Recursive Masked Sequence Tagging
- Memory and attention in deep learning
- Finite Group Equivariant Neural Networks for Games
- Relational dynamic memory networks
- A Syntax-aware Multi-task Learning Framework for Chinese Semantic Role Labeling
- Improving RNN Transducer Based ASR with Auxiliary Tasks
- Correlation Distance Skip Connection Denoising Autoencoder (CDSK-DAE) for Speech Feature Enhancement
- Recomposing the Reinforcement Learning Building Blocks with Hypernetworks
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- KGSynNet: A Novel Entity Synonyms Discovery Framework with Knowledge Graph
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- A Survey on Assessing the Generalization Envelope of Deep Neural Networks: Predictive Uncertainty, Out-of-distribution and Adversarial Samples
- Totally Deep Support Vector Machines
- The Effect of Data Ordering in Image Classification
- Seq-SetNet: Exploring Sequence Sets for Inferring Structures
- DELMU: A Deep Learning Approach to Maximising the Utility of Virtualised Millimetre-Wave Backhauls
- Towards Accurate and Compact Architectures via Neural Architecture Transformer
- DVNet: A Memory-Efficient Three-Dimensional CNN for Large-Scale Neurovascular Reconstruction
- Real-Time Limited-View CT Inpainting and Reconstruction with Dual Domain Based on Spatial Information
- Aligning Cross-Lingual Entities with Multi-Aspect Information
- MetaAugment: Sample-Aware Data Augmentation Policy Learning
- Hierarchical Multi Task Learning with Subword Contextual Embeddings for Languages with Rich Morphology
- Not All Features Are Equal: Feature Leveling Deep Neural Networks for Better Interpretation
- Deep Learning Estimation of Absorbed Dose for Nuclear Medicine Diagnostics
- A New Benchmark and Progress Toward Improved Weakly Supervised Learning
- A Graph-Based Neural Model for End-to-End Frame Semantic Parsing
- Image Annotation based on Deep Hierarchical Context Networks
- Layer Flexible Adaptive Computational Time
- Weight mechanism: adding a constant in concatenation of series connect
- Semi-tied Units for Efficient Gating in LSTM and Highway Networks
- Syntactic and Semantic-driven Learning for Open Information Extraction
- Semantic Role Labeling with Iterative Structure Refinement
- Residual networks classify inputs based on their neural transient dynamics
- Self-Teaching Networks
- Densely Connected Graph Convolutional Networks for Graph-to-Sequence Learning
- Dense Graph Convolutional Neural Networks on 3D Meshes for 3D Object Segmentation and Classification
- Improving Open Information Extraction via Iterative Rank-Aware Learning
- Boosting Network Weight Separability via Feed-Backward Reconstruction
- Text Classification based on Multiple Block Convolutional Highways
- Unbounded Output Networks for Classification
- Deep Neural Networks with Short Circuits for Improved Gradient Learning
- Sequentially Aggregated Convolutional Networks
- PR Product: A Substitute for Inner Product in Neural Networks
- Concurrently Extrapolating and Interpolating Networks for Continuous Model Generation
- Cross-lingual transfer learning for spoken language understanding
- Combining Differential Privacy and Byzantine Resilience in Distributed SGD
- What and Where to Translate: Local Mask-based Image-to-Image Translation
- Lamina-specific neuronal properties promote robust, stable signal propagation in feedforward networks
- Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
- Lattice Fusion Networks for Image Denoising
- SVGD: A Virtual Gradients Descent Method for Stochastic Optimization
- CLAR: A Cross-Lingual Argument Regularizer for Semantic Role Labeling
- Towards Better Generalization: BP-SVRG in Training Deep Neural Networks