Aggregated Residual Transformations for Deep Neural Networks
arXiv:1611.05431
Abstract
We present a simple, highly modularized network architecture for image classification. Our network is constructed by repeating a building block that aggregates a set of transformations with the same topology. Our simple design results in a homogeneous, multi-branch architecture that has only a few hyper-parameters to set. This strategy exposes a new dimension, which we call "cardinality" (the size of the set of transformations), as an essential factor in addition to the dimensions of depth and width. On the ImageNet-1K dataset, we empirically show that even under the restricted condition of maintaining complexity, increasing cardinality is able to improve classification accuracy. Moreover, increasing cardinality is more effective than going deeper or wider when we increase the capacity. Our models, named ResNeXt, are the foundations of our entry to the ILSVRC 2016 classification task in which we secured 2nd place. We further investigate ResNeXt on an ImageNet-5K set and the COCO detection set, also showing better results than its ResNet counterpart. The code and models are publicly available online.
Accepted to CVPR 2017. Code and models: https://github.com/facebookresearch/ResNeXt
References in corpus (8)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Caffe: Convolutional Architecture for Fast Feature Embedding
- WaveNet: A Generative Model for Raw Audio
- Wide Residual Networks
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Neural Machine Translation in Linear Time
- Deep Roots: Improving CNN Efficiency with Hierarchical Filter Groups
- Rigid-Motion Scattering for Texture Classification
Cited by in corpus (171)
- MLP-Mixer: An all-MLP Architecture for Vision
- Focal Loss for Dense Object Detection
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Random Erasing Data Augmentation
- AutoAugment: Learning Augmentation Policies from Data
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Born Again Neural Networks
- Advancements in Image Classification using Convolutional Neural Network
- SINet: A Scale-insensitive Convolutional Neural Network for Fast Vehicle Detection
- Channel Pruning for Accelerating Very Deep Neural Networks
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Like What You Like: Knowledge Distill via Neuron Selectivity Transfer
- Shake-Shake regularization
- Residual Attention Network for Image Classification
- A Review of Object Detection Models based on Convolutional Neural Network
- CosFace: Large Margin Cosine Loss for Deep Face Recognition
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Hierarchical Representations for Efficient Architecture Search
- Natural Adversarial Examples
- Deformable ConvNets v2: More Deformable, Better Results
- Pythia v0.1: the Winning Entry to the VQA Challenge 2018
- The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Weighted Transformer Network for Machine Translation
- Context Encoding for Semantic Segmentation
- Margin Sample Mining Loss: A Deep Learning Based Method for Person Re-identification
- XNOR-Net++: Improved Binary Neural Networks
- Using Trusted Data to Train Deep Networks on Labels Corrupted by Severe Noise
- Single-Shot Refinement Neural Network for Object Detection
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- Multi-Objective Matrix Normalization for Fine-grained Visual Recognition
- Rethinking ImageNet Pre-training
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Re-ID done right: towards good practices for person re-identification
- Measuring the Algorithmic Efficiency of Neural Networks
- Peephole: Predicting Network Performance Before Training
- Dissecting the Graphcore IPU Architecture via Microbenchmarking
- Interleaved Group Convolutions for Deep Neural Networks
- Nonlinear Approximation via Compositions
- Testing Robustness Against Unforeseen Adversaries
- PU-Net: Point Cloud Upsampling Network
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Image Captioning for Effective Use of Language Models in Knowledge-Based Visual Question Answering
- Learning Feature Pyramids for Human Pose Estimation
- Fast and Accurate Single Image Super-Resolution via Information Distillation Network
- TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris
- Analysis and Optimization of Convolutional Neural Network Architectures
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- ExFuse: Enhancing Feature Fusion for Semantic Segmentation
- Restructuring Batch Normalization to Accelerate CNN Training
- Adversarial Dropout Regularization
- Syn2Real: A New Benchmark forSynthetic-to-Real Visual Domain Adaptation
- Crystal Loss and Quality Pooling for Unconstrained Face Verification and Recognition
- Data Distillation: Towards Omni-Supervised Learning
- Shift: A Zero FLOP, Zero Parameter Alternative to Spatial Convolutions
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- BlockDrop: Dynamic Inference Paths in Residual Networks
- Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
- Towards High Performance Video Object Detection for Mobiles
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- A Good Practice Towards Top Performance of Face Recognition: Transferred Deep Feature Fusion
- MegDet: A Large Mini-Batch Object Detector
- The Herbarium Challenge 2019 Dataset
- SupportNet: solving catastrophic forgetting in class incremental learning with support data
- SCUT-FBP5500: A Diverse Benchmark Dataset for Multi-Paradigm Facial Beauty Prediction
- SUPPNet: Neural network for stellar spectrum normalisation
- A Performance Comparison of Loss Functions for Deep Face Recognition
- Quantization for Rapid Deployment of Deep Neural Networks
- NAS-FCOS: Fast Neural Architecture Search for Object Detection
- Question-Guided Hybrid Convolution for Visual Question Answering
- Practical Block-wise Neural Network Architecture Generation
- VOC-ReID: Vehicle Re-identification based on Vehicle-Orientation-Camera
- Deep Pyramidal Residual Networks with Separated Stochastic Depth
- Defense against Adversarial Attacks Using High-Level Representation Guided Denoiser
- CovidExpert: A Triplet Siamese Neural Network framework for the detection of COVID-19
- The iNaturalist Species Classification and Detection Dataset
- Graph-Based Global Reasoning Networks
- Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Dividing and Conquering Cross-Modal Recipe Retrieval: from Nearest Neighbours Baselines to SoTA
- DeepRadiologyNet: Radiologist Level Pathology Detection in CT Head Images
- BitPruning: Learning Bitlengths for Aggressive and Accurate Quantization
- End-to-end Video-level Representation Learning for Action Recognition
- Using Machine Learning at Scale in HPC Simulations with SmartSim: An Application to Ocean Climate Modeling
- Reciprocal Attention Fusion for Visual Question Answering
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Decorrelated Adversarial Learning for Age-Invariant Face Recognition
- VrR-VG: Refocusing Visually-Relevant Relationships
- Dual Path Networks for Multi-Person Human Pose Estimation
- Efficient Deep Neural Networks
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Image segmentation via Cellular Automata
- Efficient Palm-Line Segmentation with U-Net Context Fusion Module
- Large Scale Multimodal Classification Using an Ensemble of Transformer Models and Co-Attention
- Making EfficientNet More Efficient: Exploring Batch-Independent Normalization, Group Convolutions and Reduced Resolution Training
- Deep Regionlets for Object Detection
- Between-class Learning for Image Classification
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- 2nd Place Solution for Waymo Open Dataset Challenge -- 2D Object Detection
- Improved Mutual Mean-Teaching for Unsupervised Domain Adaptive Re-ID
- Towards High Performance Video Object Detection
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Rethink ReLU to Training Better CNNs
- Binarizing MobileNet via Evolution-based Searching
- NeXtVLAD: An Efficient Neural Network to Aggregate Frame-level Features for Large-scale Video Classification
- A flexible FPGA accelerator for convolutional neural networks
- Matrix and tensor decompositions for training binary neural networks
- Fibro-CoSANet: Pulmonary Fibrosis Prognosis Prediction using a Convolutional Self Attention Network
- Attention Mechanisms for Object Recognition with Event-Based Cameras
- Learning Effective Visual Relationship Detector on 1 GPU
- Detecting and counting tiny faces
- MosAIc: Finding Artistic Connections across Culture with Conditional Image Retrieval
- Why You Should Try the Real Data for the Scene Text Recognition
- Learning Dense Stereo Matching for Digital Surface Models from Satellite Imagery
- Modularity Matters: Learning Invariant Relational Reasoning Tasks
- CNN Model & Tuning for Global Road Damage Detection
- Defending Against Image Corruptions Through Adversarial Augmentations
- Automated Inline Analysis of Myocardial Perfusion MRI with Deep Learning
- ResNetX: a more disordered and deeper network architecture
- Deep Dual Pyramid Network for Barcode Segmentation using Barcode-30k Database
- Learning Less-Overlapping Representations
- Smooth Inter-layer Propagation of Stabilized Neural Networks for Classification
- Ensemble Transfer Learning for Emergency Landing Field Identification on Moderate Resource Heterogeneous Kubernetes Cluster
- Spatially-Adaptive Filter Units for Deep Neural Networks
- Towards Precise Pruning Points Detection using Semantic-Instance-Aware Plant Models for Grapevine Winter Pruning Automation
- Exploring Neural Networks Quantization via Layer-Wise Quantization Analysis
- GM-Net: Learning Features with More Efficiency
- Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
- Supervised Deep Sparse Coding Networks
- Exploring Weight Symmetry in Deep Neural Networks
- Normalization Before Shaking Toward Learning Symmetrically Distributed Representation Without Margin in Speech Emotion Recognition
- Grounded Video Description
- CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
- Are wider nets better given the same number of parameters?
- Class-Wise Difficulty-Balanced Loss for Solving Class-Imbalance
- Cross-Model Image Annotation Platform with Active Learning
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- SSAN: Separable Self-Attention Network for Video Representation Learning
- Going Deeper with Lean Point Networks
- Wasserstein Introspective Neural Networks
- C-DLinkNet: considering multi-level semantic features for human parsing
- Optimizing Block-Sparse Matrix Multiplications on CUDA with TVM
- Augmentation Inside the Network
- LogAvgExp Provides a Principled and Performant Global Pooling Operator
- FactorizeNet: Progressive Depth Factorization for Efficient Network Architecture Exploration Under Quantization Constraints
- FA-RPN: Floating Region Proposals for Face Detection
- Classifying Textual Data with Pre-trained Vision Models through Transfer Learning and Data Transformations
- Accurate and Efficient Similarity Search for Large Scale Face Recognition
- CondenseNet V2: Sparse Feature Reactivation for Deep Networks
- Image-based Vehicle Re-identification Model with Adaptive Attention Modules and Metadata Re-ranking
- MVMD: A Multi-View Approach for Enhanced Mirror Detection
- Block-Cyclic Stochastic Coordinate Descent for Deep Neural Networks
- Decision Propagation Networks for Image Classification
- Rethinking Radiology: An Analysis of Different Approaches to BraTS
- PosNeg-Balanced Anchors with Aligned Features for Single-Shot Object Detection
- Clustering and Classification Networks
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Deep Competitive Pathway Networks
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification
- Large-scale mammography CAD with Deformable Conv-Nets
- clcNet: Improving the Efficiency of Convolutional Neural Network using Channel Local Convolutions
- Compact retail shelf segmentation for mobile deployment
- Now that I can see, I can improve: Enabling data-driven finetuning of CNNs on the edge
- 1st Place Solutions for UG2+ Challenge 2021 -- (Semi-)supervised Face detection in the low light condition
- SelectScale: Mining More Patterns from Images via Selective and Soft Dropout
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio