Improving neural networks by preventing co-adaptation of feature detectors
arXiv:1207.0580
Abstract
When a large feedforward neural network is trained on a small training set, it typically performs poorly on held-out test data. This "overfitting" is greatly reduced by randomly omitting half of the feature detectors on each training case. This prevents complex co-adaptations in which a feature detector is only helpful in the context of several other specific feature detectors. Instead, each neuron learns to detect a feature that is generally helpful for producing the correct answer given the combinatorially large variety of internal contexts in which it must operate. Random "dropout" gives big improvements on many benchmark tasks and sets new records for speech and object recognition.
Cited by in corpus (159)
- Deep Learning in Neural Networks: An Overview
- Distilling the Knowledge in a Neural Network
- Conditional Generative Adversarial Nets
- How transferable are features in deep neural networks?
- Learning Transferable Features with Deep Adaptation Networks
- Striving for Simplicity: The All Convolutional Net
- Attention-Based Models for Speech Recognition
- Learning Face Representation from Scratch
- Understanding Neural Networks Through Deep Visualization
- Weight Uncertainty in Neural Networks
- Convolutional Neural Networks for Sentence Classification
- Deep Learning with Limited Numerical Precision
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- DeepID3: Face Recognition with Very Deep Neural Networks
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Deep Bayesian Active Learning with Image Data
- Training Deep Neural Networks on Noisy Labels with Bootstrapping
- Stochastic Pooling for Regularization of Deep Convolutional Neural Networks
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- On Using Monolingual Corpora in Neural Machine Translation
- A Convolutional Neural Network for Modelling Sentences
- Scalable Bayesian Optimization Using Deep Neural Networks
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Residual Networks of Residual Networks: Multilevel Residual Networks
- Generative Moment Matching Networks
- Learning with Pseudo-Ensembles
- Learning Activation Functions to Improve Deep Neural Networks
- Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Making brain-machine interfaces robust to future neural variability
- Star-galaxy Classification Using Deep Convolutional Neural Networks
- Effective Use of Word Order for Text Categorization with Convolutional Neural Networks
- Spatially-sparse convolutional neural networks
- Jet Flavor Classification in High-Energy Physics with Deep Neural Networks
- Winner-Take-All Autoencoders
- Less-forgetting Learning in Deep Neural Networks
- Multi-task Neural Networks for QSAR Predictions
- DART: Dropouts meet Multiple Additive Regression Trees
- Memory Bounded Deep Convolutional Networks
- Deep Networks with Internal Selective Attention through Feedback Connections
- Freeze-Thaw Bayesian Optimization
- Dropout Inference in Bayesian Neural Networks with Alpha-divergences
- APAC: Augmented PAttern Classification with Neural Networks
- Scale-Invariant Convolutional Neural Networks
- Enhanced Higgs to Searches with Deep Learning
- Recover Canonical-View Faces in the Wild with Deep Neural Networks
- Approaching the Computational Color Constancy as a Classification Problem through Deep Learning
- Deep Learning for Medical Image Segmentation
- Learning Fine-grained Image Similarity with Deep Ranking
- The Power of Sparsity in Convolutional Neural Networks
- Toxicity Prediction using Deep Learning
- Deep Structured Output Learning for Unconstrained Text Recognition
- Multi-Target Regression via Random Linear Target Combinations
- On the Origin of Deep Learning
- Fisher Vectors Derived from Hybrid Gaussian-Laplacian Mixture Models for Image Annotation
- LFADS - Latent Factor Analysis via Dynamical Systems
- Analyzing noise in autoencoders and deep networks
- Dataset and Neural Recurrent Sequence Labeling Model for Open-Domain Factoid Question Answering
- Deep Learning with Low Precision by Half-wave Gaussian Quantization
- Mixture of Counting CNNs: Adaptive Integration of CNNs Specialized to Specific Appearance for Crowd Counting
- A Probabilistic Theory of Deep Learning
- Neural Machine Translation with Reconstruction
- Raiders of the Lost Architecture: Kernels for Bayesian Optimization in Conditional Parameter Spaces
- Big Neural Networks Waste Capacity
- Visual Causal Feature Learning
- Heterogeneous Multi-task Learning for Human Pose Estimation with Deep Convolutional Neural Network
- Deep Temporal Appearance-Geometry Network for Facial Expression Recognition
- EmoNets: Multimodal deep learning approaches for emotion recognition in video
- When Face Recognition Meets with Deep Learning: an Evaluation of Convolutional Neural Networks for Face Recognition
- Bayesian Neural Networks for Genetic Association Studies of Complex Disease
- Techniques for Learning Binary Stochastic Feedforward Neural Networks
- On the Robustness of Convolutional Neural Networks to Internal Architecture and Weight Perturbations
- Learning to Discover Efficient Mathematical Identities
- A Bayesian encourages dropout
- Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
- Discriminative Recurrent Sparse Auto-Encoders
- On Vectorization of Deep Convolutional Neural Networks for Vision Tasks
- Deep Models for Engagement Assessment With Scarce Label Information
- Online Linear Optimization via Smoothing
- Efficient Object Localization Using Convolutional Networks
- Is Joint Training Better for Deep Auto-Encoders?
- To Drop or Not to Drop: Robustness, Consistency and Differential Privacy Properties of Dropout
- Exponentially Increasing the Capacity-to-Computation Ratio for Conditional Computation in Deep Learning
- DropNeuron: Simplifying the Structure of Deep Neural Networks
- Joint Training of Deep Boltzmann Machines
- Answer Sequence Learning with Neural Networks for Answer Selection in Community Question Answering
- Interactive Attention for Neural Machine Translation
- Deep Regression for Face Alignment
- Hierarchical learning for DNN-based acoustic scene classification
- High Performance Offline Handwritten Chinese Character Recognition Using GoogLeNet and Directional Feature Maps
- Altitude Training: Strong Bounds for Single-Layer Dropout
- Unsupervised Deep Haar Scattering on Graphs
- Learning unbiased features
- Comparing Rule-Based and Deep Learning Models for Patient Phenotyping
- Multi-path Convolutional Neural Networks for Complex Image Classification
- Neural Decision Trees
- SimNets: A Generalization of Convolutional Networks
- Object Recognition Using Deep Neural Networks: A Survey
- End-to-end Convolutional Network for Saliency Prediction
- A PCA-Based Convolutional Network
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
- Half-CNN: A General Framework for Whole-Image Regression
- Treelogy: A Novel Tree Classifier Utilizing Deep and Hand-crafted Representations
- Exploiting Sentence and Context Representations in Deep Neural Models for Spoken Language Understanding
- Understanding Locally Competitive Networks
- Dropout Training for Support Vector Machines
- Combining the Best of Graphical Models and ConvNets for Semantic Segmentation
- Variational Inference with Hamiltonian Monte Carlo
- Dropout with Expectation-linear Regularization
- Task-Oriented Learning of Word Embeddings for Semantic Relation Classification
- Teaching Deep Convolutional Neural Networks to Play Go
- Generic Object Detection With Dense Neural Patterns and Regionlets
- Abstract Learning via Demodulation in a Deep Neural Network
- Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic Oracles
- Temporal Autoencoding Restricted Boltzmann Machine
- Improved Deep Convolutional Neural Network For Online Handwritten Chinese Character Recognition using Domain-Specific Knowledge
- Character-level Chinese Writer Identification using Path Signature Feature, DropStroke and Deep CNN
- Deep Learning the Indus Script
- Deep Recurrent Neural Network for Mobile Human Activity Recognition with High Throughput
- Character-Level Language Modeling with Hierarchical Recurrent Neural Networks
- CIFAR-10: KNN-based Ensemble of Classifiers
- Word and Document Embeddings based on Neural Network Approaches
- Dependency Sensitive Convolutional Neural Networks for Modeling Sentences and Documents
- A Scale Mixture Perspective of Multiplicative Noise in Neural Networks
- HEp-2 Cell Image Classification with Deep Convolutional Neural Networks
- Scalable and Incremental Learning of Gaussian Mixture Models
- Piecewise Linear Multilayer Perceptrons and Dropout
- Exploiting the Statistics of Learning and Inference
- Improving Neural Network Generalization by Combining Parallel Circuits with Dropout
- DropSample: A New Training Method to Enhance Deep Convolutional Neural Networks for Large-Scale Unconstrained Handwritten Chinese Character Recognition
- Thoughts on a Recursive Classifier Graph: a Multiclass Network for Deep Object Recognition
- Neural Network Regularization via Robust Weight Factorization
- Autoencoder Regularized Network For Driving Style Representation Learning
- -Fields: Neural Network Nearest Neighbor Fields for Image Transforms
- On the Equivalence Between Deep NADE and Generative Stochastic Networks
- Learning Compact Convolutional Neural Networks with Nested Dropout
- Maxmin convolutional neural networks for image classification
- Can deep learning help you find the perfect match?
- Learning Machines Implemented on Non-Deterministic Hardware
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- Variational inference of latent state sequences using Recurrent Networks
- GSNs : Generative Stochastic Networks
- Recycle deep features for better object detection
- Image aesthetic evaluation using paralleled deep convolution neural network
- Regularizing Recurrent Networks - On Injected Noise and Norm-based Methods
- Lexical Translation Model Using a Deep Neural Network Architecture
- Deep Collaborative Learning for Visual Recognition
- A Multichannel Convolutional Neural Network For Cross-language Dialog State Tracking
- DropRegion Training of Inception Font Network for High-Performance Chinese Font Recognition
- Sparse, guided feature connections in an Abstract Deep Network
- Horn: A System for Parallel Training and Regularizing of Large-Scale Neural Networks
- Belief Flows of Robust Online Learning
- Search Intelligence: Deep Learning For Dominant Category Prediction
- Recognition Confidence Analysis of Handwritten Chinese Character with CNN
- Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
- Deep Transform: Time-Domain Audio Error Correction via Probabilistic Re-Synthesis
- A Distributed Deep Representation Learning Model for Big Image Data Classification
- The Long-Short Story of Movie Description
- Use Generalized Representations, But Do Not Forget Surface Features