FractalNet: Ultra-Deep Neural Networks without Residuals
arXiv:1605.07648
Abstract
We introduce a design strategy for neural network macro-architecture based on self-similarity. Repeated application of a simple expansion rule generates deep networks whose structural layouts are precisely truncated fractals. These networks contain interacting subpaths of different lengths, but do not include any pass-through or residual connections; every internal signal is transformed by a filter and nonlinearity before being seen by subsequent layers. In experiments, fractal networks match the excellent performance of standard residual networks on both CIFAR and ImageNet classification tasks, thereby demonstrating that residual representations may not be fundamental to the success of extremely deep convolutional neural networks. Rather, the key may be the ability to transition, during training, from effectively shallow to deep. We note similarities with student-teacher behavior and develop drop-path, a natural extension of dropout, to regularize co-adaptation of subpaths in fractal architectures. Such regularization allows extraction of high-performance fixed-depth subnetworks. Additionally, fractal networks exhibit an anytime property: shallow subnetworks provide a quick answer, while deeper subnetworks, with higher latency, provide a more accurate answer.
updated with ImageNet results; published as a conference paper at ICLR 2017; project page at http://people.cs.uchicago.edu/~larsson/fractalnet/
References in corpus (10)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improving neural networks by preventing co-adaptation of feature detectors
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Densely Connected Convolutional Networks
- Wide Residual Networks
- Resnet in Resnet: Generalizing Residual Architectures
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Scalable Bayesian Optimization Using Deep Neural Networks
- Fractional Max-Pooling
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
Cited by in corpus (115)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- Neural Architecture Search with Reinforcement Learning
- Densely Connected Convolutional Networks
- Transformer in Transformer
- Regularized Deep Networks in Intelligent Transportation Systems: A Taxonomy and a Case Study
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Residual Networks of Residual Networks: Multilevel Residual Networks
- Centralized Feature Pyramid for Object Detection
- Systematic evaluation of CNN advances on the ImageNet
- Shake-Shake regularization
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- GridMask Data Augmentation
- Hierarchical Representations for Efficient Architecture Search
- A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions
- A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU
- MgNet: A Unified Framework of Multigrid and Convolutional Neural Network
- Recent Advances in Deep Learning: An Overview
- Review: Deep Learning in Electron Microscopy
- Data augmentation instead of explicit regularization
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Selective Kernel Networks
- Deep Learning for Steganalysis of Diverse Data Types: A review of methods, taxonomy, challenges and future directions
- Neural Models for Information Retrieval
- How to train your MAML
- Interleaved Group Convolutions for Deep Neural Networks
- Multi-View 3D Object Detection Network for Autonomous Driving
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- ASAP: Architecture Search, Anneal and Prune
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Graph Neural Ordinary Differential Equations
- FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction
- Convolutional Neural Networks with Gated Recurrent Connections
- Deep Convolutional Neural Network Design Patterns
- Dense and Diverse Capsule Networks: Making the Capsules Learn Better
- You Only Search Once: Single Shot Neural Architecture Search via Direct Sparse Optimization
- Understanding the Disharmony between Dropout and Batch Normalization by Variance Shift
- On Robustness of Neural Ordinary Differential Equations
- NAIS-Net: Stable Deep Networks from Non-Autonomous Differential Equations
- PolyNet: A Pursuit of Structural Diversity in Very Deep Networks
- Cut-Thumbnail: A Novel Data Augmentation for Convolutional Neural Network
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Deep Back-Projection Networks for Single Image Super-resolution
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Deep Pyramidal Residual Networks
- IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
- Continuous-in-Depth Neural Networks
- Deep Back-Projection Networks For Super-Resolution
- GradAug: A New Regularization Method for Deep Neural Networks
- Rethinking Feature Distribution for Loss Functions in Image Classification
- NeuPDE: Neural Network Based Ordinary and Partial Differential Equations for Modeling Time-Dependent Data
- Review On Deep Learning Technique For Underwater Object Detection
- Automatically Evolving CNN Architectures Based on Blocks
- XNAS: Neural Architecture Search with Expert Advice
- SambaMixer: State of Health Prediction of Li-ion Batteries using Mamba State Space Models
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource Constraints
- DPDnet: A Robust People Detector using Deep Learning with an Overhead Depth Camera
- ResNet or DenseNet? Introducing Dense Shortcuts to ResNet
- Finding Better Topologies for Deep Convolutional Neural Networks by Evolution
- MSD: Multi-Self-Distillation Learning via Multi-classifiers within Deep Neural Networks
- Online Deep Learning from Doubly-Streaming Data
- MEAL: Multi-Model Ensemble via Adversarial Learning
- ThreshNet: An Efficient DenseNet Using Threshold Mechanism to Reduce Connections
- Sampled Training and Node Inheritance for Fast Evolutionary Neural Architecture Search
- 1D Convolutional Neural Network Models for Sleep Arousal Detection
- Prune and Replace NAS
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- FAS-UNet: A Novel FAS-driven Unet to Learn Variational Image Segmentation
- gSwin: Gated MLP Vision Model with Hierarchical Structure of Shifted Window
- Beyond Fine Tuning: A Modular Approach to Learning on Small Data
- Towards Learning Affine-Invariant Representations via Data-Efficient CNNs
- Scheduled DropHead: A Regularization Method for Transformer Models
- Neural Architecture Search by Estimation of Network Structure Distributions
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Weight Pruning via Adaptive Sparsity Loss
- Particle Swarm Optimisation for Evolving Deep Neural Networks for Image Classification by Evolving and Stacking Transferable Blocks
- Real-Time Steganalysis for Stream Media Based on Multi-channel Convolutional Sliding Windows
- Deblending Overlapping Galaxies in DECaLS Using Transformer-Based Algorithm: A Method Combining Multiple Bands and Data Types
- Factorized Bilinear Models for Image Recognition
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- Optimizing Neural Architecture Search using Limited GPU Time in a Dynamic Search Space: A Gene Expression Programming Approach
- Bonsai-Net: One-Shot Neural Architecture Search via Differentiable Pruners
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- Exploring Feature Reuse in DenseNet Architectures
- How Does Supernet Help in Neural Architecture Search?
- Sensor Fusion for Robot Control through Deep Reinforcement Learning
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- GM-Net: Learning Features with More Efficiency
- Genetic Network Architecture Search
- Meta-Solver for Neural Ordinary Differential Equations
- Fine-tuning Handwriting Recognition systems with Temporal Dropout
- Beyond Dropout: Feature Map Distortion to Regularize Deep Neural Networks
- Constrained Linear Data-feature Mapping for Image Classification
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- DDU-Nets: Distributed Dense Model for 3D MRI Brain Tumor Segmentation
- Orthogonal and Idempotent Transformations for Learning Deep Neural Networks
- SoFAr: Shortcut-based Fractal Architectures for Binary Convolutional Neural Networks
- Modeling Artistic Workflows for Image Generation and Editing
- RBUE: A ReLU-Based Uncertainty Estimation Method of Deep Neural Networks
- Augmented Shortcuts for Vision Transformers
- PSDNet and DPDNet: Efficient channel expansion, Depthwise-Pointwise-Depthwise Inverted Bottleneck Block
- Tandem Blocks in Deep Convolutional Neural Networks
- Convergence of backpropagation with momentum for network architectures with skip connections
- RotationOut as a Regularization Method for Neural Network
- Multi-scale Convolution Aggregation and Stochastic Feature Reuse for DenseNets
- Deep Competitive Pathway Networks
- Architectural Resilience to Foreground-and-Background Adversarial Noise
- Itsy Bitsy SpiderNet: Fully Connected Residual Network for Fraud Detection
- ConAM: Confidence Attention Module for Convolutional Neural Networks
- Adaptive Low-Rank Regularization with Damping Sequences to Restrict Lazy Weights in Deep Networks
- Dense Graph Convolutional Neural Networks on 3D Meshes for 3D Object Segmentation and Classification
- Mitigating deep double descent by concatenating inputs
- Towards Natural Robustness Against Adversarial Examples
- Pyramidal RoR for Image Classification