Wide Residual Networks
arXiv:1605.07146
Abstract
Deep residual networks were shown to be able to scale up to thousands of layers and still have improving performance. However, each fraction of a percent of improved accuracy costs nearly doubling the number of layers, and so training very deep residual networks has a problem of diminishing feature reuse, which makes these networks very slow to train. To tackle these problems, in this paper we conduct a detailed experimental study on the architecture of ResNet blocks, based on which we propose a novel architecture where we decrease depth and increase width of residual networks. We call the resulting network structures wide residual networks (WRNs) and show that these are far superior over their commonly used thin and very deep counterparts. For example, we demonstrate that even a simple 16-layer-deep wide residual network outperforms in accuracy and efficiency all previous deep residual networks, including thousand-layer-deep networks, achieving new state-of-the-art results on CIFAR, SVHN, COCO, and significant improvements on ImageNet. Our code and models are available at https://github.com/szagoruyko/wide-residual-networks
References in corpus (1)
Cited by in corpus (918)
- Denoising Diffusion Probabilistic Models
- Neural Architecture Search with Reinforcement Learning
- Bootstrap your own latent: A new approach to self-supervised Learning
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Densely Connected Convolutional Networks
- On Calibration of Modern Neural Networks
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Unsupervised Data Augmentation for Consistency Training
- Scaling Laws for Neural Language Models
- Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: the LUNA16 challenge
- Score-Based Generative Modeling through Stochastic Differential Equations
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Interpolation Consistency Training for Semi-Supervised Learning
- Random Erasing Data Augmentation
- AutoAugment: Learning Augmentation Policies from Data
- A Convergence Theory for Deep Learning via Over-Parameterization
- Embedding Watermarks into Deep Neural Networks
- Deep Learning in Video Multi-Object Tracking: A Survey
- Enhancing The Reliability of Out-of-distribution Image Detection in Neural Networks
- FractalNet: Ultra-Deep Neural Networks without Residuals
- Don't Decay the Learning Rate, Increase the Batch Size
- MentorNet: Learning Data-Driven Curriculum for Very Deep Neural Networks on Corrupted Labels
- Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
- Feature Pyramid Networks for Object Detection
- Hyper-Parameter Optimization: A Review of Algorithms and Applications
- Energy-based Out-of-distribution Detection
- The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches
- Born Again Neural Networks
- Gabor Convolutional Networks
- Residual Networks of Residual Networks: Multilevel Residual Networks
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Deep Anomaly Detection Using Geometric Transformations
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Aggregated Residual Transformations for Deep Neural Networks
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- DeepViT: Towards Deeper Vision Transformer
- CBAM: Convolutional Block Attention Module
- Bayesian Compression for Deep Learning
- Systematic evaluation of CNN advances on the ImageNet
- Residual Attention Network for Image Classification
- Generating Adversarial Examples with Adversarial Networks
- GridMask Data Augmentation
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Averaging Weights Leads to Wider Optima and Better Generalization
- On the Reconstruction of Face Images from Deep Face Templates
- Hierarchical Representations for Efficient Architecture Search
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Revisiting ResNets: Improved Training and Scaling Strategies
- Adversarial Weight Perturbation Helps Robust Generalization
- Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
- Neural Optimizer Search with Reinforcement Learning
- A Comprehensive Survey of Neural Architecture Search: Challenges and Solutions
- HeteroFL: Computation and Communication Efficient Federated Learning for Heterogeneous Clients
- You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Sparse Networks from Scratch: Faster Training without Losing Performance
- WRPN: Wide Reduced-Precision Networks
- RedNet: Residual Encoder-Decoder Network for indoor RGB-D Semantic Segmentation
- Selective Classification for Deep Neural Networks
- Automatic Skin Lesion Analysis using Large-scale Dermoscopy Images and Deep Residual Networks
- Towards the Systematic Reporting of the Energy and Carbon Footprints of Machine Learning
- Are Labels Required for Improving Adversarial Robustness?
- Improving Variational Inference with Inverse Autoregressive Flow
- A Simple Baseline for Bayesian Uncertainty in Deep Learning
- Bringing AI To Edge: From Deep Learning's Perspective
- Deep Complex Networks
- SpinalNet: Deep Neural Network with Gradual Input
- Invertible Residual Networks
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep Learning
- Learning Sparse Neural Networks through Regularization
- AdaShare: Learning What To Share For Efficient Deep Multi-Task Learning
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Context Encoding for Semantic Segmentation
- Galaxy Morphology Classification with Deep Convolutional Neural Networks
- Scaled-YOLOv4: Scaling Cross Stage Partial Network
- Optimize TSK Fuzzy Systems for Classification Problems: Mini-Batch Gradient Descent with Uniform Regularization and Batch Normalization
- Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
- Scaling provable adversarial defenses
- Knowledge Distillation in Deep Learning and its Applications
- VoxResNet: Deep Voxelwise Residual Networks for Volumetric Brain Segmentation
- Gated-SCNN: Gated Shape CNNs for Semantic Segmentation
- RobustBench: a standardized adversarial robustness benchmark
- Fixup Initialization: Residual Learning Without Normalization
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- Streaming convolutional neural networks for end-to-end learning with multi-megapixel images
- Wavelet Convolutional Neural Networks
- Varifocal-Net: A Chromosome Classification Approach using Deep Convolutional Networks
- Student-Teacher Feature Pyramid Matching for Anomaly Detection
- FMix: Enhancing Mixed Sample Data Augmentation
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance Awareness
- From Variational to Deterministic Autoencoders
- Denoising Diffusion Implicit Models
- Shallow-Deep Networks: Understanding and Mitigating Network Overthinking
- Deep Convolutional Networks as shallow Gaussian Processes
- Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Training Quantized Nets: A Deeper Understanding
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Improving Robustness Without Sacrificing Accuracy with Patch Gaussian Augmentation
- MMA Training: Direct Input Space Margin Maximization through Adversarial Training
- Deep -Means: Re-Training and Parameter Sharing with Harder Cluster Assignments for Compressing Deep Convolutions
- Selection via Proxy: Efficient Data Selection for Deep Learning
- Dissecting Neural ODEs
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised Learning
- Self-Adaptive Training: beyond Empirical Risk Minimization
- Efficient Facial Representations for Age, Gender and Identity Recognition in Organizing Photo Albums using Multi-output CNN
- What Do Compressed Deep Neural Networks Forget?
- On the Origin of Deep Learning
- Dynamic Model Pruning with Feedback
- Paraphrasing Complex Network: Network Compression via Factor Transfer
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- Measuring the Algorithmic Efficiency of Neural Networks
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks
- Attacks Which Do Not Kill Training Make Adversarial Learning Stronger
- Automatic Liver Lesion Detection using Cascaded Deep Residual Networks
- Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One
- Adversarial Self-Supervised Contrastive Learning
- An overview of mixing augmentation methods and augmentation strategies
- Regularizing CNNs with Locally Constrained Decorrelations
- Structured Denoising Diffusion Models in Discrete State-Spaces
- Discovering Parametric Activation Functions
- Generative Adversarial Residual Pairwise Networks for One Shot Learning
- Empirical Bayes Transductive Meta-Learning with Synthetic Gradients
- Deep Polynomial Neural Networks
- Heterogeneous Multilayer Generalized Operational Perceptron
- Mathematics of Deep Learning
- Interleaved Group Convolutions for Deep Neural Networks
- Nonlinear Approximation via Compositions
- Feature-map-level Online Adversarial Knowledge Distillation
- Regularizing Deep Neural Networks by Noise: Its Interpretation and Optimization
- Sequential vessel segmentation via deep channel attention network
- Scattering Networks for Hybrid Representation Learning
- Selfie: Self-supervised Pretraining for Image Embedding
- Pitfalls of In-Domain Uncertainty Estimation and Ensembling in Deep Learning
- Attentive Recurrent Comparators
- FreezeOut: Accelerate Training by Progressively Freezing Layers
- PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
- DeepCloak: Masking Deep Neural Network Models for Robustness Against Adversarial Samples
- Learning Sparse Networks Using Targeted Dropout
- Robust Pre-Training by Adversarial Contrastive Learning
- Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Exploring the Space of Black-box Attacks on Deep Neural Networks
- Vision Transformer based COVID-19 Detection using Chest X-rays
- Generalizing Convolutional Neural Networks for Equivariance to Lie Groups on Arbitrary Continuous Data
- Cloud and Cloud Shadow Segmentation for Remote Sensing Imagery via Filtered Jaccard Loss Function and Parametric Augmentation
- NBDT: Neural-Backed Decision Trees
- Residual Convolutional CTC Networks for Automatic Speech Recognition
- Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution
- Contrastive Representation Distillation
- SpaceNet: Make Free Space For Continual Learning
- Selective Feature Connection Mechanism: Concatenating Multi-layer CNN Features with a Feature Selector
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction
- ConTNet: Why not use convolution and transformer at the same time?
- Deep Learning for Environmentally Robust Speech Recognition: An Overview of Recent Developments
- BOIL: Towards Representation Change for Few-shot Learning
- Learning Feature Pyramids for Human Pose Estimation
- Analysis and Optimization of Convolutional Neural Network Architectures
- RCCNet: An Efficient Convolutional Neural Network for Histological Routine Colon Cancer Nuclei Classification
- i-RevNet: Deep Invertible Networks
- Neural Tangents: Fast and Easy Infinite Neural Networks in Python
- Comparison of Batch Normalization and Weight Normalization Algorithms for the Large-scale Image Classification
- Detection of Face Recognition Adversarial Attacks
- Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures
- Augmented Neural ODEs
- Multiscale Deep Equilibrium Models
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- Deep Predictive Coding Network for Object Recognition
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- A Constructive Prediction of the Generalization Error Across Scales
- Accurate Pulmonary Nodule Detection in Computed Tomography Images Using Deep Convolutional Neural Networks
- Augment your batch: better training with larger batches
- The Early Phase of Neural Network Training
- Auto-Ensemble: An Adaptive Learning Rate Scheduling based Deep Learning Model Ensembling
- Three Mechanisms of Weight Decay Regularization
- Quadratic Suffices for Over-parametrization via Matrix Chernoff Bound
- Pruning Neural Networks at Initialization: Why are We Missing the Mark?
- Deep Convolutional Neural Network Design Patterns
- Structured Consistency Loss for semi-supervised semantic segmentation
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- Detecting Dementia from Speech and Transcripts using Transformers
- Data Valuation using Reinforcement Learning
- TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism
- Overfitting in adversarially robust deep learning
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Scale-Equivariant Steerable Networks
- Probabilistic Binary Neural Networks
- Greedy Policy Search: A Simple Baseline for Learnable Test-Time Augmentation
- DeepMarks: A Digital Fingerprinting Framework for Deep Neural Networks
- BlackMarks: Blackbox Multibit Watermarking for Deep Neural Networks
- Learning perturbation sets for robust machine learning
- ABC: Auxiliary Balanced Classifier for Class-imbalanced Semi-supervised Learning
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- Understanding the Disharmony between Dropout and Batch Normalization by Variance Shift
- Multi-Residual Networks: Improving the Speed and Accuracy of Residual Networks
- Graph-based Knowledge Distillation by Multi-head Attention Network
- MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling
- More is Less: A More Complicated Network with Less Inference Complexity
- Class-Imbalanced Semi-Supervised Learning
- SemiFL: Semi-Supervised Federated Learning for Unlabeled Clients with Alternate Training
- PolyNet: A Pursuit of Structural Diversity in Very Deep Networks
- Bridging the Accuracy Gap for 2-bit Quantized Neural Networks (QNN)
- Unsupervised Outlier Detection using Memory and Contrastive Learning
- Generalization Error of Invariant Classifiers
- Comparative evaluation of CNN architectures for Image Caption Generation
- Accelerating Deep Learning by Focusing on the Biggest Losers
- Generalized ODIN: Detecting Out-of-distribution Image without Learning from Out-of-distribution Data
- Variational Information Distillation for Knowledge Transfer
- Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
- Do Wider Neural Networks Really Help Adversarial Robustness?
- Scaling of neural-network quantum states for time evolution
- Deep Gamblers: Learning to Abstain with Portfolio Theory
- GFF: Gated Fully Fusion for Semantic Segmentation
- Calibration of Deep Probabilistic Models with Decoupled Bayesian Neural Networks
- Defensive Quantization: When Efficiency Meets Robustness
- A Deep Learning Approach for Pose Estimation from Volumetric OCT Data
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- Large batch size training of neural networks with adversarial training and second-order information
- Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative Models
- Zero-shot Knowledge Transfer via Adversarial Belief Matching
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Deep Layer Aggregation
- Oriented Response Networks
- Towards Real-Time Multi-Object Tracking
- Multi-Task Zipping via Layer-wise Neuron Sharing
- Leveraging the Feature Distribution in Transfer-based Few-Shot Learning
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Hyperparameter Ensembles for Robustness and Uncertainty Quantification
- Contrastive Model Inversion for Data-Free Knowledge Distillation
- Feature Fusion for Online Mutual Knowledge Distillation
- Tensor Normalization and Full Distribution Training
- Geometry-aware Instance-reweighted Adversarial Training
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Label Embedding Network: Learning Label Representation for Soft Training of Deep Networks
- CryptoNAS: Private Inference on a ReLU Budget
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- FINN-R: An End-to-End Deep-Learning Framework for Fast Exploration of Quantized Neural Networks
- Sharing Residual Units Through Collective Tensor Factorization in Deep Neural Networks
- Open-set Label Noise Can Improve Robustness Against Inherent Label Noise
- Zero Time Waste: Recycling Predictions in Early Exit Neural Networks
- Lost in Pruning: The Effects of Pruning Neural Networks beyond Test Accuracy
- Learning to Defend by Learning to Attack
- Improving Auto-Augment via Augmentation-Wise Weight Sharing
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform
- Deep Pyramidal Residual Networks
- A Good Practice Towards Top Performance of Face Recognition: Transferred Deep Feature Fusion
- Towards Understanding Label Smoothing
- Assessing Post-Disaster Damage from Satellite Imagery using Semi-Supervised Learning Techniques
- Privado: Practical and Secure DNN Inference with Enclaves
- Distillation Early Stopping? Harvesting Dark Knowledge Utilizing Anisotropic Information Retrieval For Overparameterized Neural Network
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- Wide Compression: Tensor Ring Nets
- A Theoretical Framework for Robustness of (Deep) Classifiers against Adversarial Examples
- On Feature Normalization and Data Augmentation
- Self-Supervised Learning For Few-Shot Image Classification
- FlexConv: Continuous Kernel Convolutions with Differentiable Kernel Sizes
- Automated Human Cell Classification in Sparse Datasets using Few-Shot Learning
- Training DNNs with Hybrid Block Floating Point
- Scaling Laws for Deep Learning
- Interpolated Adversarial Training: Achieving Robust Neural Networks without Sacrificing Too Much Accuracy
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- Training Shallow and Thin Networks for Acceleration via Knowledge Distillation with Conditional Adversarial Networks
- Continuous-in-Depth Neural Networks
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study
- Vision Permutator: A Permutable MLP-Like Architecture for Visual Recognition
- S-Net: A Scalable Convolutional Neural Network for JPEG Compression Artifact Reduction
- Robust Neural Networks using Randomized Adversarial Training
- Neural Networks with Recurrent Generative Feedback
- Consistent Estimators for Learning to Defer to an Expert
- Rethinking Feature Distribution for Loss Functions in Image Classification
- When Ensembling Smaller Models is More Efficient than Single Large Models
- GradAug: A New Regularization Method for Deep Neural Networks
- Greedy Layerwise Learning Can Scale to ImageNet
- Two Sides of the Same Coin: White-box and Black-box Attacks for Transfer Learning
- Towards Understanding Fast Adversarial Training
- DPGN: Distribution Propagation Graph Network for Few-shot Learning
- DeepReDuce: ReLU Reduction for Fast Private Inference
- On the Validity of Bayesian Neural Networks for Uncertainty Estimation
- Provable Filter Pruning for Efficient Neural Networks
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- FEED: Feature-level Ensemble for Knowledge Distillation
- OpenMatch: Open-set Consistency Regularization for Semi-supervised Learning with Outliers
- Mitigating Bias in Calibration Error Estimation
- Adaptive Regularization of Labels
- Test-time Batch Statistics Calibration for Covariate Shift
- Sill-Net: Feature Augmentation with Separated Illumination Representation
- Hybrid Discriminative-Generative Training via Contrastive Learning
- End-to-End Multi-Task Learning with Attention
- Towards Effective Low-bitwidth Convolutional Neural Networks
- Towards Compact and Robust Deep Neural Networks
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- Lightweight Residual Densely Connected Convolutional Neural Network
- Practical Block-wise Neural Network Architecture Generation
- How Does Mixup Help With Robustness and Generalization?
- FLOPs as a Direct Optimization Objective for Learning Sparse Neural Networks
- Preserving Earlier Knowledge in Continual Learning with the Help of All Previous Feature Extractors
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free
- Towards Feature Space Adversarial Attack
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource Constraints
- VFlow: More Expressive Generative Flows with Variational Data Augmentation
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning
- Reliable Adversarial Distillation with Unreliable Teachers
- Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
- Learning Chained Deep Features and Classifiers for Cascade in Object Detection
- The Evolution of Out-of-Distribution Robustness Throughout Fine-Tuning
- Regularizing Activation Distribution for Training Binarized Deep Networks
- SirenAttack: Generating Adversarial Audio for End-to-End Acoustic Systems
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural Networks
- Improving Semantic Segmentation via Decoupled Body and Edge Supervision
- Differentiable Model Compression via Pseudo Quantization Noise
- A Simple Fine-tuning Is All You Need: Towards Robust Deep Learning Via Adversarial Fine-tuning
- Parle: parallelizing stochastic gradient descent
- One Model to Rule them all: Multitask and Multilingual Modelling for Lexical Analysis
- Do deep nets really need weight decay and dropout?
- Normalized Direction-preserving Adam
- Scaling the Scattering Transform: Deep Hybrid Networks
- Learning to Learn Variational Semantic Memory
- BlockQNN: Efficient Block-wise Neural Network Architecture Generation
- HCM: Hardware-Aware Complexity Metric for Neural Network Architectures
- Deep Learning based approach to detect Customer Age, Gender and Expression in Surveillance Video
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
- What causes the test error? Going beyond bias-variance via ANOVA
- Projection Convolutional Neural Networks for 1-bit CNNs via Discrete Back Propagation
- Finding Better Topologies for Deep Convolutional Neural Networks by Evolution
- Towards an Adversarially Robust Normalization Approach
- Few-Shot Image Recognition by Predicting Parameters from Activations
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- Large Scale Visual Food Recognition
- Don't Just Blame Over-parametrization for Over-confidence: Theoretical Analysis of Calibration in Binary Classification
- Adversarial Training Versus Weight Decay
- Attended Temperature Scaling: A Practical Approach for Calibrating Deep Neural Networks
- Deep learning for photoacoustic imaging: a survey
- Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
- MixNorm: Test-Time Adaptation Through Online Normalization Estimation
- When Semi-Supervised Learning Meets Transfer Learning: Training Strategies, Models and Datasets
- The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
- PowerNorm: Rethinking Batch Normalization in Transformers
- Adjusting for Dropout Variance in Batch Normalization and Weight Initialization
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge Distillation
- Post-Training Piecewise Linear Quantization for Deep Neural Networks
- Interpreting Deep Classifier by Visual Distillation of Dark Knowledge
- Co-Mixup: Saliency Guided Joint Mixup with Supermodular Diversity
- Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
- Learning Energy-Based Models by Diffusion Recovery Likelihood
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Deep Speech Enhancement for Reverberated and Noisy Signals using Wide Residual Networks
- Understanding the Interaction of Adversarial Training with Noisy Labels
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- Recognizing Birds from Sound - The 2018 BirdCLEF Baseline System
- Efficient Neural Network Training via Forward and Backward Propagation Sparsification
- Learning Neural Network Subspaces
- Confidence Scoring Using Whitebox Meta-models with Linear Classifier Probes
- Deep Learning Assessment of galaxy morphology in S-PLUS DataRelease 1
- Multi-scale Adaptive Task Attention Network for Few-Shot Learning
- Are Visual Explanations Useful? A Case Study in Model-in-the-Loop Prediction
- Rethinking Re-Sampling in Imbalanced Semi-Supervised Learning
- A New Semi-supervised Learning Benchmark for Classifying View and Diagnosing Aortic Stenosis from Echocardiograms
- K-Hairstyle: A Large-scale Korean Hairstyle Dataset for Virtual Hair Editing and Hairstyle Classification
- C-RPNs: Promoting Object Detection in real world via a Cascade Structure of Region Proposal Networks
- Convolution Neural Network Architecture Learning for Remote Sensing Scene Classification
- Advancing System Performance with Redundancy: From Biological to Artificial Designs
- Can We Gain More from Orthogonality Regularizations in Training Deep CNNs?
- Generalization in Machine Learning via Analytical Learning Theory
- Learning Anytime Predictions in Neural Networks via Adaptive Loss Balancing
- Densely Guided Knowledge Distillation using Multiple Teacher Assistants
- Large Scale Private Learning via Low-rank Reparametrization
- Rethinking Bottleneck Structure for Efficient Mobile Network Design
- Beyond Class-Conditional Assumption: A Primary Attempt to Combat Instance-Dependent Label Noise
- OVANet: One-vs-All Network for Universal Domain Adaptation
- Progressive Feature Fusion Network for Realistic Image Dehazing
- Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation
- Constructing Deep Neural Networks by Bayesian Network Structure Learning
- Active Convolution: Learning the Shape of Convolution for Image Classification
- E-swish: Adjusting Activations to Different Network Depths
- Age Group and Gender Estimation in the Wild with Deep RoR Architecture
- Taylorized Training: Towards Better Approximation of Neural Network Training at Finite Width
- Shaping representations through communication: community size effect in artificial learning systems
- Trained Rank Pruning for Efficient Deep Neural Networks
- WeMix: How to Better Utilize Data Augmentation
- An Enhanced Convolutional Neural Network in Side-Channel Attacks and Its Visualization
- Real-time Universal Style Transfer on High-resolution Images via Zero-channel Pruning
- Using Machine Learning at Scale in HPC Simulations with SmartSim: An Application to Ocean Climate Modeling
- LocalDrop: A Hybrid Regularization for Deep Neural Networks
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks
- Distributional Generalization: A New Kind of Generalization
- Temperature check: theory and practice for training models with softmax-cross-entropy losses
- Few-shot Classification via Adaptive Attention
- RIGA: Covert and Robust White-Box Watermarking of Deep Neural Networks
- Learning to play the Chess Variant Crazyhouse above World Champion Level with Deep Neural Networks and Human Data
- Active Image Synthesis for Efficient Labeling
- On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- Large Margin Deep Networks for Classification
- Sample-level CNN Architectures for Music Auto-tagging Using Raw Waveforms
- Online Adversarial Purification based on Self-Supervision
- Multiresolution Convolutional Autoencoders
- Time for a Background Check! Uncovering the impact of Background Features on Deep Neural Networks
- Learning data augmentation policies using augmented random search
- Deep Metric Transfer for Label Propagation with Limited Annotated Data
- Domain-specific Communication Optimization for Distributed DNN Training
- Test time Adaptation through Perturbation Robustness
- DNN or k-NN: That is the Generalize vs. Memorize Question
- On the design of convolutional neural networks for automatic detection of Alzheimer's disease
- On the Relation between Color Image Denoising and Classification
- The Dual Information Bottleneck
- Bridging the Performance Gap between FGSM and PGD Adversarial Training
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Convolutional Residual Memory Networks
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Orthogonalizing Convolutional Layers with the Cayley Transform
- Unsupervised Deep Features for Remote Sensing Image Matching via Discriminator Network
- ATHENA: A Framework based on Diverse Weak Defenses for Building Adversarial Defense
- Understanding Generalization in Deep Learning via Tensor Methods
- Diversity inducing Information Bottleneck in Model Ensembles
- Contextual Dropout: An Efficient Sample-Dependent Dropout Module
- LTD: Low Temperature Distillation for Gradient Masking-free Adversarial Training
- Low-Precision Batch-Normalized Activations
- Fully Decoupled Neural Network Learning Using Delayed Gradients
- On Improving Adversarial Transferability of Vision Transformers
- Understanding Generalization in Adversarial Training via the Bias-Variance Decomposition
- Sparse Weight Activation Training
- Neural Architecture Dilation for Adversarial Robustness
- Full deep neural network training on a pruned weight budget
- A Survey on Deep Domain Adaptation and Tiny Object Detection Challenges, Techniques and Datasets
- Rethinking FUN: Frequency-Domain Utilization Networks
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- Efficient Adversarial Training with Transferable Adversarial Examples
- Rethinking and Improving the Robustness of Image Style Transfer
- Sanity Checks for Lottery Tickets: Does Your Winning Ticket Really Win the Jackpot?
- Cyclic Differentiable Architecture Search
- Deep Q-Networks for Accelerating the Training of Deep Neural Networks
- Batch Group Normalization
- Sampled Training and Node Inheritance for Fast Evolutionary Neural Architecture Search
- Are all outliers alike? On Understanding the Diversity of Outliers for Detecting OODs
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
- Effective Training of Convolutional Neural Networks with Low-bitwidth Weights and Activations
- Deep Neural Networks with Multi-Branch Architectures Are Less Non-Convex
- Towards Practical Lottery Ticket Hypothesis for Adversarial Training
- Anomalous Example Detection in Deep Learning: A Survey
- Convolutional Networks with Dense Connectivity
- CHOPT : Automated Hyperparameter Optimization Framework for Cloud-Based Machine Learning Platforms
- Transfer learning based few-shot classification using optimal transport mapping from preprocessed latent space of backbone neural network
- Should Ensemble Members Be Calibrated?
- Query Attack via Opposite-Direction Feature:Towards Robust Image Retrieval
- Adversarial Robustness Assessment: Why both and Attacks Are Necessary
- Stochastic Security: Adversarial Defense Using Long-Run Dynamics of Energy-Based Models
- Combining Ensembles and Data Augmentation can Harm your Calibration
- Compact Global Descriptor for Neural Networks
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Online Hyper-parameter Learning for Auto-Augmentation Strategy
- A Gap-Based Framework for Chinese Word Segmentation via Very Deep Convolutional Networks
- Revisiting One-vs-All Classifiers for Predictive Uncertainty and Out-of-Distribution Detection in Neural Networks
- Adversarial Token Attacks on Vision Transformers
- Intra-model Variability in COVID-19 Classification Using Chest X-ray Images
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Robust or Private? Adversarial Training Makes Models More Vulnerable to Privacy Attacks
- On the Necessity and Effectiveness of Learning the Prior of Variational Auto-Encoder
- Adversarially Adaptive Normalization for Single Domain Generalization
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- A Critical Evaluation of Open-World Machine Learning
- I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively
- Weakly Supervised Clustering by Exploiting Unique Class Count
- Boosting Black-Box Attack with Partially Transferred Conditional Adversarial Distribution
- Encoder Fusion Network with Co-Attention Embedding for Referring Image Segmentation
- IRLAS: Inverse Reinforcement Learning for Architecture Search
- Adaptive Consistency Regularization for Semi-Supervised Transfer Learning
- PENNI: Pruned Kernel Sharing for Efficient CNN Inference
- Learning Better Internal Structure of Words for Sequence Labeling
- Is Support Set Diversity Necessary for Meta-Learning?
- Orthogonal Deep Neural Networks
- A Generic Network Compression Framework for Sequential Recommender Systems
- Deep CNNs for Peripheral Blood Cell Classification
- Graph-based Interpolation of Feature Vectors for Accurate Few-Shot Classification
- On the asymptotics of wide networks with polynomial activations
- flexgrid2vec: Learning Efficient Visual Representations Vectors
- Adversarial Visual Robustness by Causal Intervention
- MetaDistiller: Network Self-Boosting via Meta-Learned Top-Down Distillation
- Coresets for Robust Training of Neural Networks against Noisy Labels
- Faster Neural Network Training with Approximate Tensor Operations
- Automated Learning Rate Scheduler for Large-batch Training
- Deep Transform and Metric Learning Network: Wedding Deep Dictionary Learning and Neural Networks
- Revisiting the Evaluation of Uncertainty Estimation and Its Application to Explore Model Complexity-Uncertainty Trade-Off
- Improving Query Efficiency of Black-box Adversarial Attack
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- S3NAS: Fast NPU-aware Neural Architecture Search Methodology
- Detecting Overfitting via Adversarial Examples
- Context-aware stacked convolutional neural networks for classification of breast carcinomas in whole-slide histopathology images
- CIFAR-10: KNN-based Ensemble of Classifiers
- Truth or Backpropaganda? An Empirical Investigation of Deep Learning Theory
- Random Noise Defense Against Query-Based Black-Box Attacks
- On the interplay between data structure and loss function in classification problems
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- OpenCoS: Contrastive Semi-supervised Learning for Handling Open-set Unlabeled Data
- RayS: A Ray Searching Method for Hard-label Adversarial Attack
- The State of Knowledge Distillation for Classification
- For self-supervised learning, Rationality implies generalization, provably
- No MCMC for me: Amortized sampling for fast and stable training of energy-based models
- EEGdenoiseNet: A benchmark dataset for end-to-end deep learning solutions of EEG denoising
- Towards Learning Affine-Invariant Representations via Data-Efficient CNNs
- Implicit Rugosity Regularization via Data Augmentation
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Rethinking Recurrent Neural Networks and Other Improvements for Image Classification
- ResNet After All? Neural ODEs and Their Numerical Solution
- Learnable Descent Algorithm for Nonsmooth Nonconvex Image Reconstruction
- A Loss Curvature Perspective on Training Instability in Deep Learning
- Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models
- Guided Interpolation for Adversarial Training
- Convolution Aware Initialization
- Strength in Numbers: Trading-off Robustness and Computation via Adversarially-Trained Ensembles
- Generalized Adaptation for Few-Shot Learning
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- DeepOBS: A Deep Learning Optimizer Benchmark Suite
- Enhancing Adversarial Defense by k-Winners-Take-All
- Computation-Efficient Knowledge Distillation via Uncertainty-Aware Mixup
- Rethinking Curriculum Learning with Incremental Labels and Adaptive Compensation
- Rethinking the Number of Channels for the Convolutional Neural Network
- A Survey on Fault-tolerance in Distributed Optimization and Machine Learning
- SOAR: Second-Order Adversarial Regularization
- Self-Adaptive Training: Bridging Supervised and Self-Supervised Learning
- Parallel Grid Pooling for Data Augmentation
- Listen carefully and tell: an audio captioning system based on residual learning and gammatone audio representation
- Harmonic Networks: Integrating Spectral Information into CNNs
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Countering Noisy Labels By Learning From Auxiliary Clean Labels
- TinyAction Challenge: Recognizing Real-world Low-resolution Activities in Videos
- Trade-offs and Guarantees of Adversarial Representation Learning for Information Obfuscation
- An Asymptotically Optimal Multi-Armed Bandit Algorithm and Hyperparameter Optimization
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Reversed Active Learning based Atrous DenseNet for Pathological Image Classification
- Multi-head Knowledge Distillation for Model Compression
- Soft Labeling Affects Out-of-Distribution Detection of Deep Neural Networks
- Regional Adversarial Training for Better Robust Generalization
- Revisiting Adversarial Robustness Distillation: Robust Soft Labels Make Student Better
- Zero and Few Shot Learning with Semantic Feature Synthesis and Competitive Learning
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- Auxiliary Task Update Decomposition: The Good, The Bad and The Neutral
- Towards Oracle Knowledge Distillation with Neural Architecture Search
- Explainable Deep Learning for Uncovering Actionable Scientific Insights for Materials Discovery and Design
- Negative sampling in semi-supervised learning
- Generalizing Deep Models for Overhead Image Segmentation Through Getis-Ord Gi* Pooling
- SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)
- Chameleon: Learning Model Initializations Across Tasks With Different Schemas
- Robust and On-the-fly Dataset Denoising for Image Classification
- Why bigger is not always better: on finite and infinite neural networks
- Particle Swarm Optimisation for Evolving Deep Neural Networks for Image Classification by Evolving and Stacking Transferable Blocks
- On Detecting GANs and Retouching based Synthetic Alterations
- DeepPeep: Exploiting Design Ramifications to Decipher the Architecture of Compact DNNs
- Rethink ReLU to Training Better CNNs
- MetaDelta: A Meta-Learning System for Few-shot Image Classification
- Achieving Real-Time LiDAR 3D Object Detection on a Mobile Device
- Residual Convolutional Neural Network Revisited with Active Weighted Mapping
- Pufferfish: Communication-efficient Models At No Extra Cost
- How and When Adversarial Robustness Transfers in Knowledge Distillation?
- RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs
- An Experience-based Direct Generation approach to Automatic Image Cropping
- Drop-Activation: Implicit Parameter Reduction and Harmonic Regularization
- Supervised COSMOS Autoencoder: Learning Beyond the Euclidean Loss!
- S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima
- Simulating Personal Food Consumption Patterns using a Modified Markov Chain
- Overcoming Catastrophic Forgetting by Generative Regularization
- Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
- Intra-Ensemble in Neural Networks
- QEBA: Query-Efficient Boundary-Based Blackbox Attack
- Locally Enhanced Self-Attention: Combining Self-Attention and Convolution as Local and Context Terms
- M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition
- Modular Universal Reparameterization: Deep Multi-task Learning Across Diverse Domains
- Learning Optimal Data Augmentation Policies via Bayesian Optimization for Image Classification Tasks
- Context-aware PolyUNet for Liver and Lesion Segmentation from Abdominal CT Images
- Semi-Supervised and Active Few-Shot Learning with Prototypical Networks
- Supervised Deep Neural Networks (DNNs) for Pricing/Calibration of Vanilla/Exotic Options Under Various Different Processes
- Few-Shot Learning by Integrating Spatial and Frequency Representation
- Lottery Jackpots Exist in Pre-trained Models
- Understanding the Impact of Label Granularity on CNN-based Image Classification
- Functional Gradient Boosting based on Residual Network Perception
- -SGD: Optimizing ReLU Neural Networks in its Positively Scale-Invariant Space
- 3D Aggregated Faster R-CNN for General Lesion Detection
- Semi-Supervised Learning Enabled by Multiscale Deep Neural Network Inversion
- Data Quality Matters For Adversarial Training: An Empirical Study
- Factorized Bilinear Models for Image Recognition
- An Empirical Study and Analysis on Open-Set Semi-Supervised Learning
- Online Knowledge Distillation with Diverse Peers
- EDAS: Efficient and Differentiable Architecture Search
- Does Data Augmentation Benefit from Split BatchNorms
- GBCNs: Genetic Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs
- A Simple Dynamic Learning Rate Tuning Algorithm For Automated Training of DNNs
- ARiA: Utilizing Richard's Curve for Controlling the Non-monotonicity of the Activation Function in Deep Neural Nets
- RankingMatch: Delving into Semi-Supervised Learning with Consistency Regularization and Ranking Loss
- Knowledge distillation for optimization of quantized deep neural networks
- Learning Representations that Support Robust Transfer of Predictors
- One-bit Supervision for Image Classification
- Switching Transferable Gradient Directions for Query-Efficient Black-Box Adversarial Attacks
- Stochastic Gradient Descent with Hyperbolic-Tangent Decay on Classification
- DelugeNets: Deep Networks with Efficient and Flexible Cross-layer Information Inflows
- Deep Dimension Reduction for Supervised Representation Learning
- Calibration of Neural Networks using Splines
- Energy-efficient Amortized Inference with Cascaded Deep Classifiers
- Neural Network Encapsulation
- Improving Adversarial Robustness via Unlabeled Out-of-Domain Data
- Dynamic Sparse Graph for Efficient Deep Learning
- A Study of Checkpointing in Large Scale Training of Deep Neural Networks
- OnlineAugment: Online Data Augmentation with Less Domain Knowledge
- DAAS: Differentiable Architecture and Augmentation Policy Search
- MetaInfoNet: Learning Task-Guided Information for Sample Reweighting
- On Calibration of Mixup Training for Deep Neural Networks
- Curriculum By Smoothing
- Toward Adversarial Robustness via Semi-supervised Robust Training
- Hidden-Fold Networks: Random Recurrent Residuals Using Sparse Supermasks
- SORT: Second-Order Response Transform for Visual Recognition
- PydMobileNet: Improved Version of MobileNets with Pyramid Depthwise Separable Convolution
- DeepNNK: Explaining deep models and their generalization using polytope interpolation
- Semantically-Conditioned Negative Samples for Efficient Contrastive Learning
- Simplified Stochastic Feedforward Neural Networks
- Lifelong Learning with Sketched Structural Regularization
- A New Look at Ghost Normalization
- Learning Deep Representations Using Convolutional Auto-encoders with Symmetric Skip Connections
- DeepSearch: A Simple and Effective Blackbox Attack for Deep Neural Networks
- VACL: Variance-Aware Cross-Layer Regularization for Pruning Deep Residual Networks
- TandemNet: Distilling Knowledge from Medical Images Using Diagnostic Reports as Optional Semantic References
- ACDC: Weight Sharing in Atom-Coefficient Decomposed Convolution
- MDCN: Multi-Scale, Deep Inception Convolutional Neural Networks for Efficient Object Detection
- StackMix: A complementary Mix algorithm
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- Characterizing Membership Privacy in Stochastic Gradient Langevin Dynamics
- Representation Transfer by Optimal Transport
- Parameter Prediction for Unseen Deep Architectures
- Training convolutional neural networks with cheap convolutions and online distillation
- Heuristic Rank Selection with Progressively Searching Tensor Ring Network
- TRADI: Tracking deep neural network weight distributions for uncertainty estimation
- Student Network Learning via Evolutionary Knowledge Distillation
- Backpropagation for Implicit Spectral Densities
- Single Layer Predictive Normalized Maximum Likelihood for Out-of-Distribution Detection
- Binarized Neural Architecture Search
- Improving Model Robustness with Latent Distribution Locally and Globally
- How Out-of-Distribution Data Hurts Semi-Supervised Learning
- Improve Adversarial Robustness via Weight Penalization on Classification Layer
- Adversarial Transformations for Semi-Supervised Learning
- Putting visual object recognition in context
- Semi-supervised Learning by Latent Space Energy-Based Model of Symbol-Vector Coupling
- Classification of simulated radio signals using Wide Residual Networks for use in the search for extra-terrestrial intelligence
- ResNetX: a more disordered and deeper network architecture
- Contextual Classification Using Self-Supervised Auxiliary Models for Deep Neural Networks
- Compressing gradients by exploiting temporal correlation in momentum-SGD
- Unsupervised Temperature Scaling: An Unsupervised Post-Processing Calibration Method of Deep Networks
- Learning Less-Overlapping Representations
- Full-attention based Neural Architecture Search using Context Auto-regression
- Robust Visual Object Tracking with Two-Stream Residual Convolutional Networks
- Self-Orthogonality Module: A Network Architecture Plug-in for Learning Orthogonal Filters
- NLNL: Negative Learning for Noisy Labels
- Weighting Is Worth the Wait: Bayesian Optimization with Importance Sampling
- Leveraging Class Similarity to Improve Deep Neural Network Robustness
- Learnable Adaptive Cosine Estimator (LACE) for Image Classification
- Non-Parametric Transformation Networks
- Improved Adversarial Training via Learned Optimizer
- Decoder Choice Network for Meta-Learning
- Gradually Updated Neural Networks for Large-Scale Image Recognition
- Building a Regular Decision Boundary with Deep Networks
- Wavelet Denoised-ResNet CNN and LightGBM Method to Predict Forex Rate of Change
- Robustness, Privacy, and Generalization of Adversarial Training
- Dueling Decoders: Regularizing Variational Autoencoder Latent Spaces
- Adaptive Feature Alignment for Adversarial Training
- Implicit bias of deep linear networks in the large learning rate phase
- Exploring Covariate and Concept Shift for Detection and Calibration of Out-of-Distribution Data
- Optimizing Neural Architecture Search using Limited GPU Time in a Dynamic Search Space: A Gene Expression Programming Approach
- Large-Scale Optimal Transport via Adversarial Training with Cycle-Consistency
- Medical Knowledge-Guided Deep Learning for Imbalanced Medical Image Classification
- Compressing Heavy-Tailed Weight Matrices for Non-Vacuous Generalization Bounds
- Compressive Visual Representations
- Ex uno plures: Splitting One Model into an Ensemble of Subnetworks
- Meta-Cal: Well-controlled Post-hoc Calibration by Ranking
- On Out-of-distribution Detection with Energy-based Models
- Smoothness Analysis of Adversarial Training
- A Unified Game-Theoretic Interpretation of Adversarial Robustness
- Follow Your Path: a Progressive Method for Knowledge Distillation
- Attention-based fusion of semantic boundary and non-boundary information to improve semantic segmentation
- Confidence Calibration with Bounded Error Using Transformations
- Improving Anytime Prediction with Parallel Cascaded Networks and a Temporal-Difference Loss
- Progressive Representative Labeling for Deep Semi-Supervised Learning
- An Attention Module for Convolutional Neural Networks
- Why Layer-Wise Learning is Hard to Scale-up and a Possible Solution via Accelerated Downsampling
- Mixtures of Laplace Approximations for Improved Post-Hoc Uncertainty in Deep Learning
- MetaPerturb: Transferable Regularizer for Heterogeneous Tasks and Architectures
- Improving Classifier Confidence using Lossy Label-Invariant Transformations
- Uniform Priors for Data-Efficient Transfer
- SuperNet -- An efficient method of neural networks ensembling
- A Novel Learnable Gradient Descent Type Algorithm for Non-convex Non-smooth Inverse Problems
- Utilizing Network Properties to Detect Erroneous Inputs
- Regularized Evolutionary Population-Based Training
- Multi-objective Search of Robust Neural Architectures against Multiple Types of Adversarial Attacks
- Data augmentation with Mobius transformations
- Perceptually Constrained Adversarial Attacks
- Random Bundle: Brain Metastases Segmentation Ensembling through Annotation Randomization
- Defective Convolutional Networks
- Image Restoration Using Deep Regulated Convolutional Networks
- Metric learning by Similarity Network for Deep Semi-Supervised Learning
- Exploring Model Robustness with Adaptive Networks and Improved Adversarial Training
- Pseudo-Representation Labeling Semi-Supervised Learning
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- Ensemble Transfer Learning for Emergency Landing Field Identification on Moderate Resource Heterogeneous Kubernetes Cluster
- On the Empirical Neural Tangent Kernel of Standard Finite-Width Convolutional Neural Network Architectures
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Actionable Attribution Maps for Scientific Machine Learning
- Deep Neural Network Ensembles
- Knowledge Distillation Meets Self-Supervision
- Understanding and Diagnosing Vulnerability under Adversarial Attacks
- Learned Proximal Networks for Quantitative Susceptibility Mapping
- Extending Label Smoothing Regularization with Self-Knowledge Distillation
- Margin-Based Regularization and Selective Sampling in Deep Neural Networks
- On the Orthogonality of Knowledge Distillation with Other Techniques: From an Ensemble Perspective
- Improving Differentially Private Models with Active Learning
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- Approximated Orthonormal Normalisation in Training Neural Networks
- Selective sampling for accelerating training of deep neural networks
- A study on the role of subsidiary information in replay attack spoofing detection
- Deep Learning in Memristive Nanowire Networks
- On Norm-Agnostic Robustness of Adversarial Training
- Truncating Wide Networks using Binary Tree Architectures
- GM-Net: Learning Features with More Efficiency
- Supervised Deep Sparse Coding Networks
- CrescendoNet: A Simple Deep Convolutional Neural Network with Ensemble Behavior
- PSDNet and DPDNet: Efficient channel expansion, Depthwise-Pointwise-Depthwise Inverted Bottleneck Block
- Adaptive Clustering of Robust Semantic Representations for Adversarial Image Purification
- Complementary Relation Contrastive Distillation
- Bandwidth-based Step-Sizes for Non-Convex Stochastic Optimization
- Classification of Long Noncoding RNA Elements Using Deep Convolutional Neural Networks and Siamese Networks
- A Convergence Theory Towards Practical Over-parameterized Deep Neural Networks
- Road images augmentation with synthetic traffic signs using neural networks
- MetaAugment: Sample-Aware Data Augmentation Policy Learning
- Radius-margin bounds for deep neural networks
- Differentiable Feature Aggregation Search for Knowledge Distillation
- FOCUS: Familiar Objects in Common and Uncommon Settings
- Integrating Multiple Receptive Fields through Grouped Active Convolution
- Out-Of-Distribution Detection With Subspace Techniques And Probabilistic Modeling Of Features
- Gradient-based Data Augmentation for Semi-Supervised Learning
- SIPA: A Simple Framework for Efficient Networks
- Boosting the Performance of Semi-Supervised Learning with Unsupervised Clustering
- Parallel Deep Neural Networks Have Zero Duality Gap
- ResNEsts and DenseNEsts: Block-based DNN Models with Improved Representation Guarantees
- An Exploration of Mimic Architectures for Residual Network Based Spectral Mapping
- Fused Deep Convolutional Neural Network for Precision Diagnosis of COVID-19 Using Chest X-Ray Images
- Feature Distillation With Guided Adversarial Contrastive Learning
- Contrastive Weight Regularization for Large Minibatch SGD
- Out-of-Distribution Detection Using an Ensemble of Self Supervised Leave-out Classifiers
- On the Marginal Benefit of Active Learning: Does Self-Supervision Eat Its Cake?
- An Autonomous Approach to Measure Social Distances and Hygienic Practices during COVID-19 Pandemic in Public Open Spaces
- How Convolutional Neural Network Architecture Biases Learned Opponency and Colour Tuning
- Data Augmentation via Structured Adversarial Perturbations
- Scalable Bayesian neural networks by layer-wise input augmentation
- LogAvgExp Provides a Principled and Performant Global Pooling Operator
- Communication-Compressed Adaptive Gradient Method for Distributed Nonconvex Optimization
- What augmentations are sensitive to hyper-parameters and why?
- Signature-Graph Networks
- GCCN: Global Context Convolutional Network
- Data Summarization via Bilevel Optimization
- Building Compact and Robust Deep Neural Networks with Toeplitz Matrices
- Effective Regularization Through Loss-Function Metalearning
- Temporally Resolution Decrement: Utilizing the Shape Consistency for Higher Computational Efficiency
- FALCON: Feature Driven Selective Classification for Energy-Efficient Image Recognition
- Towards Accurate Quantization and Pruning via Data-free Knowledge Transfer
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Matching Distributions via Optimal Transport for Semi-Supervised Learning
- Content-Adaptive Pixel Discretization to Improve Model Robustness
- Augmentation Inside the Network
- Removing Undesirable Feature Contributions Using Out-of-Distribution Data
- DOI: Divergence-based Out-of-Distribution Indicators via Deep Generative Models
- Deep Collaborative Learning for Visual Recognition
- ReRankMatch: Semi-Supervised Learning with Semantics-Oriented Similarity Representation
- Sequential Random Network for Fine-grained Image Classification
- CRL: Class Representative Learning for Image Classification
- Local Patch AutoAugment with Multi-Agent Collaboration
- Learning a Discriminant Latent Space with Neural Discriminant Analysis
- On The Distribution of Penultimate Activations of Classification Networks
- SPI-Optimizer: an integral-Separated PI Controller for Stochastic Optimization
- Domain Invariant Adversarial Learning
- D-PCN: Parallel Convolutional Networks for Image Recognition via a Discriminator
- Dynamic Defense Approach for Adversarial Robustness in Deep Neural Networks via Stochastic Ensemble Smoothed Model
- CircConv: A Structured Convolution with Low Complexity
- Using Anomaly Feature Vectors for Detecting, Classifying and Warning of Outlier Adversarial Examples
- GAN-based Pose-aware Regulation for Video-based Person Re-identification
- Photozilla: A Large-Scale Photography Dataset and Visual Embedding for 20 Photography Styles
- ESFNet: Efficient Network for Building Extraction from High-Resolution Aerial Images
- Deep Asymmetric Multi-task Feature Learning
- On Efficient Uncertainty Estimation for Resource-Constrained Mobile Applications
- Convergence of backpropagation with momentum for network architectures with skip connections
- A Low-Compexity Deep Learning Framework For Acoustic Scene Classification
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- AI Driven Road Maintenance Inspection
- Greedy Network Enlarging
- Prune Your Model Before Distill It
- Representation Quality Of Neural Networks Links To Adversarial Attacks and Defences
- Multi-level Feature Fusion-based CNN for Local Climate Zone Classification from Sentinel-2 Images: Benchmark Results on the So2Sat LCZ42 Dataset
- Knowledge Transfer Graph for Deep Collaborative Learning
- Self-Denoising Neural Networks for Few Shot Learning
- Understanding the Role of Self-Supervised Learning in Out-of-Distribution Detection Task
- Training robust anomaly detection using ML-Enhanced simulations
- State-Reification Networks: Improving Generalization by Modeling the Distribution of Hidden Representations
- Noisy Feature Mixup
- End-to-End Bengali Speech Recognition
- Calibrated Adversarial Training
- Efficient Modelling Across Time of Human Actions and Interactions
- Circulant Binary Convolutional Networks: Enhancing the Performance of 1-bit DCNNs with Circulant Back Propagation
- Learning to Collaborate for User-Controlled Privacy
- Max and Coincidence Neurons in Neural Networks
- Robust Temporal Ensembling for Learning with Noisy Labels
- Modulated Self-attention Convolutional Network for VQA
- Comb Convolution for Efficient Convolutional Architecture
- Online Filter Clustering and Pruning for Efficient Convnets
- Customer Analytics using Surveillance Video
- Super Interaction Neural Network
- A Simple Riemannian Manifold Network for Image Set Classification
- Optimizing Data Augmentation Policy Through Random Unidimensional Search
- Hierarchical Auxiliary Learning
- RotationOut as a Regularization Method for Neural Network
- Multi-Scale Spatially-Asymmetric Recalibration for Image Classification
- Variance Reduction in Deep Learning: More Momentum is All You Need
- A Novel Unsupervised Post-Processing Calibration Method for DNNS with Robustness to Domain Shift
- Decoder-free Robustness Disentanglement without (Additional) Supervision
- Matching the Clinical Reality: Accurate OCT-Based Diagnosis From Few Labels
- Self-supervision of Feature Transformation for Further Improving Supervised Learning
- Mixture separability loss in a deep convolutional network for image classification
- Bayesian Optimized 1-Bit CNNs
- Deep Adaptive Wavelet Network
- ADA-Tucker: Compressing Deep Neural Networks via Adaptive Dimension Adjustment Tucker Decomposition
- Boosted CVaR Classification
- Exponential Discriminative Metric Embedding in Deep Learning
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification
- Improving Visual Recognition using Ambient Sound for Supervision
- ChebLieNet: Invariant Spectral Graph NNs Turned Equivariant by Riemannian Geometry on Lie Groups
- Learning from Matured Dumb Teacher for Fine Generalization
- Stochastic Whitening Batch Normalization
- Spectral Tensor Train Parameterization of Deep Learning Layers
- Batch Normalization and the impact of batch structure on the behavior of deep convolution networks
- Learning Multiple Categories on Deep Convolution Networks
- Constant Random Perturbations Provide Adversarial Robustness with Minimal Effect on Accuracy
- Computing Class Hierarchies from Classifiers
- Convolutional Hashing for Automated Scene Matching
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning
- Differentiable Combinatorial Losses through Generalized Gradients of Linear Programs
- Network Adjustment: Channel Search Guided by FLOPs Utilization Ratio
- ConAM: Confidence Attention Module for Convolutional Neural Networks
- Dual Head Adversarial Training
- "You might also like this model": Data Driven Approach for Recommending Deep Learning Models for Unknown Image Datasets
- CloudifierNet -- Deep Vision Models for Artificial Image Processing
- Deep Feature Pyramid Convolutional Networks with In-Place Activated Batch Normalization for Automated Skin Lesion Boundary Segmentation
- Relative Depth Order Estimation Using Multi-scale Densely Connected Convolutional Networks
- Fast Jacobian-Vector Product for Deep Networks
- Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
- Multi-Pretext Attention Network for Few-shot Learning with Self-supervision
- Learning with Hyperspherical Uniformity
- Improved Robustness of Vision Transformer via PreLayerNorm in Patch Embedding
- Selective Output Smoothing Regularization: Regularize Neural Networks by Softening Output Distributions
- Orientation Convolutional Networks for Image Recognition
- Gabor filter incorporated CNN for compression
- Privacy-Aware Activity Classification from First Person Office Videos
- Architectural Resilience to Foreground-and-Background Adversarial Noise
- Mitigating deep double descent by concatenating inputs
- Associative Convolutional Layers
- Joint Regularization on Activations and Weights for Efficient Neural Network Pruning
- Using an ensemble color space model to tackle adversarial examples
- Graphs for deep learning representations
- SegCodeNet: Color-Coded Segmentation Masks for Activity Detection from Wearable Cameras
- clcNet: Improving the Efficiency of Convolutional Neural Network using Channel Local Convolutions
- Skeptical Deep Learning with Distribution Correction
- Privacy Preserving Recalibration under Domain Shift
- Built-in Elastic Transformations for Improved Robustness
- Strategies for Robust Image Classification
- Sequenced-Replacement Sampling for Deep Learning
- FROST: Faster and more Robust One-shot Semi-supervised Training
- Effect of Various Regularizers on Model Complexities of Neural Networks in Presence of Input Noise
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- Learning Inward Scaled Hypersphere Embedding: Exploring Projections in Higher Dimensions
- FILTRA: Rethinking Steerable CNN by Filter Transform
- Multiclass non-Adversarial Image Synthesis, with Application to Classification from Very Small Sample
- Langevin Cooling for Domain Translation
- Multi-scale Convolution Aggregation and Stochastic Feature Reuse for DenseNets
- Unsupervised and Supervised Structure Learning for Protein Contact Prediction
- Generalization by Recognizing Confusion
- SelectScale: Mining More Patterns from Images via Selective and Soft Dropout
- Rethinking Uncertainty in Deep Learning: Whether and How it Improves Robustness
- Self-Gradient Networks
- Enhancing Transformation-based Defenses using a Distribution Classifier
- Exploiting Nontrivial Connectivity for Automatic Speech Recognition
- Multi-way Encoding for Robustness
- Reconciling Feature-Reuse and Overfitting in DenseNet with Specialized Dropout
- Deep Goal-Oriented Clustering
- PERMDNN: Efficient Compressed DNN Architecture with Permuted Diagonal Matrices
- KDExplainer: A Task-oriented Attention Model for Explaining Knowledge Distillation
- Semantics-Preserving Adversarial Training
- Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network Training
- Deep Virtual Networks for Memory Efficient Inference of Multiple Tasks
- Deep Competitive Pathway Networks
- Sub-clusters of Normal Data for Anomaly Detection
- Light Multi-segment Activation for Model Compression
- Neural Optimization Kernel: Towards Robust Deep Learning
- Quantal synaptic dilution enhances sparse encoding and dropout regularisation in deep networks
- Understanding Classifier Mistakes with Generative Models
- Unbounded Output Networks for Classification
- Attend and Rectify: a Gated Attention Mechanism for Fine-Grained Recovery
- Deep Anchored Convolutional Neural Networks
- Constraining Logits by Bounded Function for Adversarial Robustness
- ShuffleBlock: Shuffle to Regularize Deep Convolutional Neural Networks
- Spatially Attentive Output Layer for Image Classification
- PR Product: A Substitute for Inner Product in Neural Networks
- Aggregated Learning: A Deep Learning Framework Based on Information-Bottleneck Vector Quantization
- What can linear interpolation of neural network loss landscapes tell us?
- Learning to Reweight with Deep Interactions
- Predictive Geological Mapping with Convolution Neural Network Using Statistical Data Augmentation on a 3D Model
- Go Small and Similar: A Simple Output Decay Brings Better Performance