SGDR: Stochastic Gradient Descent with Warm Restarts
arXiv:1608.03983
Abstract
Restart techniques are common in gradient-free optimization to deal with multimodal functions. Partial warm restarts are also gaining popularity in gradient-based optimization to improve the rate of convergence in accelerated gradient schemes to deal with ill-conditioned functions. In this paper, we propose a simple warm restart technique for stochastic gradient descent to improve its anytime performance when training deep neural networks. We empirically study its performance on the CIFAR-10 and CIFAR-100 datasets, where we demonstrate new state-of-the-art results at 3.14% and 16.21%, respectively. We also demonstrate its advantages on a dataset of EEG recordings and on a downsampled version of the ImageNet dataset. Our source code is available at https://github.com/loshchil/SGDR
ICLR 2017 conference paper
References in corpus (8)
- ADADELTA: An Adaptive Learning Rate Method
- Deep learning with convolutional neural networks for EEG decoding and visualization
- Densely Connected Convolutional Networks
- Wide Residual Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Snapshot Ensembles: Train 1, get M for free
- Deep Pyramidal Residual Networks
- Adaptive Restart for Accelerated Gradient Schemes
Cited by in corpus (556)
- YOLOv4: Optimal Speed and Accuracy of Object Detection
- A Simple Framework for Contrastive Learning of Visual Representations
- Learning Transferable Visual Models From Natural Language Supervision
- Gaussian Error Linear Units (GELUs)
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- AutoAugment: Learning Augmentation Policies from Data
- Efficient Neural Architecture Search via Parameter Sharing
- Don't Decay the Learning Rate, Increase the Batch Size
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo Labeling
- Learning Spatial Fusion for Single-Shot Object Detection
- YOLOP: You Only Look Once for Panoptic Driving Perception
- Energy-based Out-of-distribution Detection
- Born Again Neural Networks
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- CogView: Mastering Text-to-Image Generation via Transformers
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- DeepViT: Towards Deeper Vision Transformer
- Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty
- Shake-Shake regularization
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Learned Step Size Quantization
- Deep Gradient Learning for Efficient Camouflaged Object Detection
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- Reducing Transformer Depth on Demand with Structured Dropout
- MCUNet: Tiny Deep Learning on IoT Devices
- High-Performance Large-Scale Image Recognition Without Normalization
- Machine-Learning-Based Diagnostics of EEG Pathology
- Contrastive Cross-site Learning with Redesigned Net for COVID-19 CT Classification
- Stand-Alone Self-Attention in Vision Models
- Decoupling Representation and Classifier for Long-Tailed Recognition
- Text2Motion: From Natural Language Instructions to Feasible Plans
- Revisiting ResNets: Improved Training and Scaling Strategies
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs
- Neural Optimizer Search with Reinforcement Learning
- Simple And Efficient Architecture Search for Convolutional Neural Networks
- The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study
- Progressive Neural Architecture Search
- Audiovisual SlowFast Networks for Video Recognition
- Discrimination-aware Network Pruning for Deep Model Compression
- Path-Level Network Transformation for Efficient Architecture Search
- Bag of Freebies for Training Object Detection Neural Networks
- DeepSOCIAL: Social Distancing Monitoring and Infection Risk Assessment in COVID-19 Pandemic
- SlowFast Networks for Video Recognition
- Explaining Neural Scaling Laws
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- Self-supervised Pretraining of Visual Features in the Wild
- Optimization for deep learning: theory and algorithms
- Banach Wasserstein GAN
- Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
- DPT: Deformable Patch-based Transformer for Visual Recognition
- Training with Quantization Noise for Extreme Model Compression
- Cyclical Stochastic Gradient MCMC for Bayesian Deep Learning
- Benchmarking Neural Network Robustness to Common Corruptions and Surface Variations
- Pelee: A Real-Time Object Detection System on Mobile Devices
- Once-for-All: Train One Network and Specialize it for Efficient Deployment
- CondenseNet: An Efficient DenseNet using Learned Group Convolutions
- Aerial Images Processing for Car Detection using Convolutional Neural Networks: Comparison between Faster R-CNN and YoloV3
- BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning
- Self-Adaptive Training: beyond Empirical Risk Minimization
- An End-to-End Earthquake Detection Method for Joint Phase Picking and Association using Deep Learning
- Towards Automated Deep Learning: Efficient Joint Neural Architecture and Hyperparameter Search
- Memory-Efficient Implementation of DenseNets
- Video Classification with Channel-Separated Convolutional Networks
- Adversarial Self-Supervised Contrastive Learning
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
- How to train your MAML
- Advancing COVID-19 Diagnosis with Privacy-Preserving Collaboration in Artificial Intelligence
- FreezeOut: Accelerate Training by Progressively Freezing Layers
- GPT-GNN: Generative Pre-Training of Graph Neural Networks
- NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm
- Self-paced and self-consistent co-training for semi-supervised image segmentation
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- BNAS:An Efficient Neural Architecture Search Approach Using Broad Scalable Architecture
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Non-intrusive reduced order modeling of natural convection in porous media using convolutional autoencoders: comparison with linear subspace techniques
- ASAP: Architecture Search, Anneal and Prune
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- Graph Neural Ordinary Differential Equations
- TimeMatch: Unsupervised Cross-Region Adaptation by Temporal Shift Estimation
- Analysis and Optimization of Convolutional Neural Network Architectures
- Long-Short Transformer: Efficient Transformers for Language and Vision
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- Multiscale Vision Transformers
- SELF: Learning to Filter Noisy Labels with Self-Ensembling
- On Warm-Starting Neural Network Training
- OSLNet: Deep Small-Sample Classification with an Orthogonal Softmax Layer
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- AlphaX: eXploring Neural Architectures with Deep Neural Networks and Monte Carlo Tree Search
- Adaptive Input Representations for Neural Language Modeling
- ktrain: A Low-Code Library for Augmented Machine Learning
- RobustART: Benchmarking Robustness on Architecture Design and Training Techniques
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal Effect
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Vertex and Energy Reconstruction in JUNO with Machine Learning Methods
- Primordial non-Gaussianity from the Completed SDSS-IV extended Baryon Oscillation Spectroscopic Survey I: Catalogue Preparation and Systematic Mitigation
- Learning Memory-guided Normality for Anomaly Detection
- Identifying Melanoma Images using EfficientNet Ensemble: Winning Solution to the SIIM-ISIC Melanoma Classification Challenge
- Semi-Supervised Neural Architecture Search
- Using convolutional neural networks to predict galaxy metallicity from three-color images
- Hierarchical ResNeXt Models for Breast Cancer Histology Image Classification
- Real-Time Semantic Segmentation via Multiply Spatial Fusion Network
- DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator Search
- An Exponential Learning Rate Schedule for Deep Learning
- Thinking in Frequency: Face Forgery Detection by Mining Frequency-aware Clues
- ReSSL: Relational Self-Supervised Learning with Weak Augmentation
- Decoding ECoG signal into 3D hand translation using deep learning
- Multi-hop Question Generation with Graph Convolutional Network
- Neural Networks Versus Conventional Filters for Inertial-Sensor-based Attitude Estimation
- BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models
- DeepCOVIDExplainer: Explainable COVID-19 Diagnosis Based on Chest X-ray Images
- Learning Loss for Test-Time Augmentation
- Class-Imbalanced Semi-Supervised Learning
- SemiFL: Semi-Supervised Federated Learning for Unlabeled Clients with Alternate Training
- Contrastive Learning with Stronger Augmentations
- Knowledge Distillation via Route Constrained Optimization
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- Compounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network
- Deeper Insights into Weight Sharing in Neural Architecture Search
- Deep Convolutional Neural Networks with Merge-and-Run Mappings
- Negative Margin Matters: Understanding Margin in Few-shot Classification
- Towards Efficient Training for Neural Network Quantization
- TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learning
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Pruning Convolutional Neural Networks with Self-Supervision
- Robust Out-of-distribution Detection for Neural Networks
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
- Automated Identification of Cell Populations in Flow Cytometry Data with Transformers
- Real or Fake? Learning to Discriminate Machine from Human Generated Text
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- Computation Reallocation for Object Detection
- MST: Masked Self-Supervised Transformer for Visual Representation
- Real-Time Likelihood-Free Inference of Roman Binary Microlensing Events with Amortized Neural Posterior Estimation
- Source Free Domain Adaptation with Image Translation
- Noisy Machines: Understanding Noisy Neural Networks and Enhancing Robustness to Analog Hardware Errors Using Distillation
- HINet: Half Instance Normalization Network for Image Restoration
- Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
- Model-based Adversarial Imitation Learning
- On Feature Normalization and Data Augmentation
- FlexConv: Continuous Kernel Convolutions with Differentiable Kernel Sizes
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- Theory-Inspired Path-Regularized Differential Network Architecture Search
- Learning image representations for anomaly detection: application to discovery of histological alterations in drug development
- Learning to solve inverse problems using Wasserstein loss
- Incremental False Negative Detection for Contrastive Learning
- GDRNPP: A Geometry-guided and Fully Learning-based Object Pose Estimator
- Pre-trained Language Model Representations for Language Generation
- Continuous conditional generative adversarial networks for data-driven solutions of poroelasticity with heterogeneous material properties
- EmotionX-KU: BERT-Max based Contextual Emotion Classifier
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- A community-powered search of machine learning strategy space to find NMR property prediction models
- Consistent Estimators for Learning to Defer to an Expert
- SemiFed: Semi-supervised Federated Learning with Consistency and Pseudo-Labeling
- Point Cloud Registration using Representative Overlapping Points
- Geometry-Consistent Neural Shape Representation with Implicit Displacement Fields
- Regularizing Deep Networks with Semantic Data Augmentation
- Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates
- Fine-Grained Visual Classification via Progressive Multi-Granularity Training of Jigsaw Patches
- XNAS: Neural Architecture Search with Expert Advice
- A Baseline for Few-Shot Image Classification
- Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint Optimization
- SIMILAR: Submodular Information Measures Based Active Learning In Realistic Scenarios
- GhostSR: Learning Ghost Features for Efficient Image Super-Resolution
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- Value Iteration Networks on Multiple Levels of Abstraction
- UAVs Beneath the Surface: Cooperative Autonomy for Subterranean Search and Rescue in DARPA SubT
- Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-Resolution
- Parle: parallelizing stochastic gradient descent
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- Neural Predictor for Neural Architecture Search
- SE-SSD: Self-Ensembling Single-Stage Object Detector From Point Cloud
- LoCo: Local Contrastive Representation Learning
- Simplifying Hamiltonian and Lagrangian Neural Networks via Explicit Constraints
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural Networks
- Normalized Direction-preserving Adam
- Investigation of Compressor Cascade Flow Using Physics- Informed Neural Networks with Adaptive Learning Strategy
- Residual Energy-Based Models for Text
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- DPP-Net: Device-aware Progressive Search for Pareto-optimal Neural Architectures
- Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need
- FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
- BlockQNN: Efficient Block-wise Neural Network Architecture Generation
- Detailed 2D-3D Joint Representation for Human-Object Interaction
- Detecting COVID-19 from digitized ECG printouts using 1D convolutional neural networks
- Geometric Back-projection Network for Point Cloud Classification
- Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
- MoViNets: Mobile Video Networks for Efficient Video Recognition
- Learning Video Representations from Textual Web Supervision
- TPLinker: Single-stage Joint Extraction of Entities and Relations Through Token Pair Linking
- GTA: Global Temporal Attention for Video Action Understanding
- Learning Dynamic Alignment via Meta-filter for Few-shot Learning
- Improving Semi-supervised Federated Learning by Reducing the Gradient Diversity of Models
- FPConv: Learning Local Flattening for Point Convolution
- Robust Learning Under Label Noise With Iterative Noise-Filtering
- On the Application of Danskin's Theorem to Derivative-Free Minimax Optimization
- Penalizing Top Performers: Conservative Loss for Semantic Segmentation Adaptation
- MixStyle Neural Networks for Domain Generalization and Adaptation
- Anytime Inference with Distilled Hierarchical Neural Ensembles
- CLAR: Contrastive Learning of Auditory Representations
- RetinaTrack: Online Single Stage Joint Detection and Tracking
- Learning Neural Network Subspaces
- Stochastic-YOLO: Efficient Probabilistic Object Detection under Dataset Shifts
- Boosting Discriminative Visual Representation Learning with Scenario-Agnostic Mixup
- AdderNet: Do We Really Need Multiplications in Deep Learning?
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- AutoSpeech: Neural Architecture Search for Speaker Recognition
- Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
- DASS: Differentiable Architecture Search for Sparse neural networks
- SimPLE: Similar Pseudo Label Exploitation for Semi-Supervised Classification
- Restricted Recurrent Neural Networks
- Gated Convolutional Networks with Hybrid Connectivity for Image Classification
- On Learning Rates and Schrödinger Operators
- EnsembleNet: End-to-End Optimization of Multi-headed Models
- Three-Stream 3D/1D CNN for Fine-Grained Action Classification and Segmentation in Table Tennis
- FaceGuard: Proactive Deepfake Detection
- ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
- MovieNet: A Holistic Dataset for Movie Understanding
- A Deep Learning-based Quality Assessment and Segmentation System with a Large-scale Benchmark Dataset for Optical Coherence Tomographic Angiography Image
- Variational Disentanglement for Domain Generalization
- Hierarchical Opacity Propagation for Image Matting
- Ultra Fast Structure-aware Deep Lane Detection
- AdCo: Adversarial Contrast for Efficient Learning of Unsupervised Representations from Self-Trained Negative Adversaries
- Learning data augmentation policies using augmented random search
- Simultaneous lesion and neuroanatomy segmentation in Multiple Sclerosis using deep neural networks
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- LineMVGNN: Anti-Money Laundering with Line-Graph-Assisted Multi-View Graph Neural Networks
- MintNet: Building Invertible Neural Networks with Masked Convolutions
- CPT: Efficient Deep Neural Network Training via Cyclic Precision
- PaStaNet: Toward Human Activity Knowledge Engine
- The Two Regimes of Deep Network Training
- GRAD-MATCH: Gradient Matching based Data Subset Selection for Efficient Deep Model Training
- Gradient-only line searches: An Alternative to Probabilistic Line Searches
- Momentum^2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
- Can weight sharing outperform random architecture search? An investigation with TuNAS
- Convolution-Weight-Distribution Assumption: Rethinking the Criteria of Channel Pruning
- AdaS: Adaptive Scheduling of Stochastic Gradients
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
- Joint gender and age estimation based on speech signals using x-vectors and transfer learning
- NDD20: A large-scale few-shot dolphin dataset for coarse and fine-grained categorisation
- Uncertainty Estimates and Multi-Hypotheses Networks for Optical Flow
- Deep Q-Networks for Accelerating the Training of Deep Neural Networks
- Recurrent Localization Networks applied to the Lippmann-Schwinger Equation
- Highly Efficient Natural Image Matting
- Cyclic Differentiable Architecture Search
- Deep Clustering by Semantic Contrastive Learning
- Neural Architecture Search using Deep Neural Networks and Monte Carlo Tree Search
- FASTER Recurrent Networks for Efficient Video Classification
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- Annealed Stein Variational Gradient Descent
- Semi-Supervised Segmentation of Salt Bodies in Seismic Images using an Ensemble of Convolutional Neural Networks
- A Distributed Hierarchical SGD Algorithm with Sparse Global Reduction
- MnasFPN: Learning Latency-aware Pyramid Architecture for Object Detection on Mobile Devices
- NTIRE 2020 Challenge on Perceptual Extreme Super-Resolution: Methods and Results
- Real-time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-driving Images
- SwishNet: A Fast Convolutional Neural Network for Speech, Music and Noise Classification and Segmentation
- Auto-PyTorch Tabular: Multi-Fidelity MetaLearning for Efficient and Robust AutoDL
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- Adversarially Adaptive Normalization for Single Domain Generalization
- PERF-Net: Pose Empowered RGB-Flow Net
- Improving Transformation Invariance in Contrastive Representation Learning
- A Contrastive Learning Approach for Training Variational Autoencoder Priors
- Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
- Delving Deep into the Generalization of Vision Transformers under Distribution Shifts
- Automated Learning Rate Scheduler for Large-batch Training
- A strong baseline for image and video quality assessment
- Temporal Context Aggregation for Video Retrieval with Contrastive Learning
- Disentangling Adaptive Gradient Methods from Learning Rates
- Automated Identification of Thoracic Pathology from Chest Radiographs with Enhanced Training Pipeline
- Adaptive Consistency Regularization for Semi-Supervised Transfer Learning
- SDWNet: A Straight Dilated Network with Wavelet Transformation for Image Deblurring
- Domain Attention Consistency for Multi-Source Domain Adaptation
- Empirical study towards understanding line search approximations for training neural networks
- Instance Localization for Self-supervised Detection Pretraining
- Automated Model Design and Benchmarking of 3D Deep Learning Models for COVID-19 Detection with Chest CT Scans
- ImmuNeCS: Neural Committee Search by an Artificial Immune System
- DCANet: Learning Connected Attentions for Convolutional Neural Networks
- Hierarchical Feature Embedding for Attribute Recognition
- Self-supervised Motion Learning from Static Images
- Confidence Adaptive Regularization for Deep Learning with Noisy Labels
- 1st Place Solutions of Waymo Open Dataset Challenge 2020 -- 2D Object Detection Track
- Gated Channel Transformation for Visual Recognition
- Feature Boosting, Suppression, and Diversification for Fine-Grained Visual Classification
- RankPose: Learning Generalised Feature with Rank Supervision for Head Pose Estimation
- Revisiting consistency for semi-supervised semantic segmentation
- Monotone-Value Neural Networks: Exploiting Preference Monotonicity in Combinatorial Assignment
- A Realistic Evaluation of Semi-Supervised Learning for Fine-Grained Classification
- Assessing Robustness of Deep learning Methods in Dermatological Workflow
- Classification of Dermoscopy Images using Deep Learning
- Relational Action Forecasting
- A Comparison of Approaches to Document-level Machine Translation
- Learning Representations for Predicting Future Activities
- Zooming SlowMo: An Efficient One-Stage Framework for Space-Time Video Super-Resolution
- Unsupervised Object-Level Representation Learning from Scene Images
- Towards Single Stage Weakly Supervised Semantic Segmentation
- Parallel Grid Pooling for Data Augmentation
- Countering Noisy Labels By Learning From Auxiliary Clean Labels
- CoL: Contrastive Continual Learning
- Relative stability toward diffeomorphisms indicates performance in deep nets
- Meta Approach to Data Augmentation Optimization
- Natural Image Matting via Guided Contextual Attention
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- Weight Pruning via Adaptive Sparsity Loss
- Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning
- Shift Invariance Can Reduce Adversarial Robustness
- Self-Supervised Learning by Estimating Twin Class Distributions
- An analysis of over-sampling labeled data in semi-supervised learning with FixMatch
- Stochastic Training is Not Necessary for Generalization
- A Variational U-Net for Weather Forecasting
- Multi-objective Neural Architecture Search via Non-stationary Policy Gradient
- Nondeterminism and Instability in Neural Network Optimization
- A Machine Learning Approach to Correcting Atmospheric Seeing in Solar Flare Observations
- ScaleNAS: One-Shot Learning of Scale-Aware Representations for Visual Recognition
- VINNAS: Variational Inference-based Neural Network Architecture Search
- Rethinking the Number of Channels for the Convolutional Neural Network
- More than Encoder: Introducing Transformer Decoder to Upsample
- Using Mode Connectivity for Loss Landscape Analysis
- Automatic Post-Stroke Lesion Segmentation on MR Images using 3D Residual Convolutional Neural Network
- Embedding Transfer with Label Relaxation for Improved Metric Learning
- STN-Homography: estimate homography parameters directly
- Backbone Can Not be Trained at Once: Rolling Back to Pre-trained Network for Person Re-Identification
- The 2ST-UNet for Pneumothorax Segmentation in Chest X-Rays using ResNet34 as a Backbone for U-Net
- Contextualizing ASR Lattice Rescoring with Hybrid Pointer Network Language Model
- Robust and On-the-fly Dataset Denoising for Image Classification
- Back-tracing Representative Points for Voting-based 3D Object Detection in Point Clouds
- SimTriplet: Simple Triplet Representation Learning with a Single GPU
- Traffic4cast 2020 -- Graph Ensemble Net and the Importance of Feature And Loss Function Design for Traffic Prediction
- Improving Accuracy of Binary Neural Networks using Unbalanced Activation Distribution
- Label-Aware Distribution Calibration for Long-tailed Classification
- Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework
- RW-Resnet: A Novel Speech Anti-Spoofing Model Using Raw Waveform
- REX: Revisiting Budgeted Training with an Improved Schedule
- Not All Memories are Created Equal: Learning to Forget by Expiring
- ResAtom System: Protein and Ligand Affinity Prediction Model Based on Deep Learning
- Busy-Quiet Video Disentangling for Video Classification
- VA-RED: Video Adaptive Redundancy Reduction
- Re-labeling ImageNet: from Single to Multi-Labels, from Global to Localized Labels
- VLG-Net: Video-Language Graph Matching Network for Video Grounding
- Towards Practical Lipreading with Distilled and Efficient Models
- Learning Architectures for Binary Networks
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-based Approach
- Repetitive Reprediction Deep Decipher for Semi-Supervised Learning
- Scribosermo: Fast Speech-to-Text models for German and other Languages
- RankingMatch: Delving into Semi-Supervised Learning with Consistency Regularization and Ranking Loss
- Better SGD using Second-order Momentum
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual Data
- A Generalizable Approach to Learning Optimizers
- BS-NAS: Broadening-and-Shrinking One-Shot NAS with Searchable Numbers of Channels
- Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error
- MOI-Mixer: Improving MLP-Mixer with Multi Order Interactions in Sequential Recommendation
- Improving Adversarial Robustness via Unlabeled Out-of-Domain Data
- Stochastic Gradient Descent with Hyperbolic-Tangent Decay on Classification
- Learning Efficient Video Representation with Video Shuffle Networks
- On Feature Decorrelation in Self-Supervised Learning
- A Simple Dynamic Learning Rate Tuning Algorithm For Automated Training of DNNs
- Temporal Modulation Network for Controllable Space-Time Video Super-Resolution
- Group Equivariant Stand-Alone Self-Attention For Vision
- memeBot: Towards Automatic Image Meme Generation
- PONAS: Progressive One-shot Neural Architecture Search for Very Efficient Deployment
- Neural Architecture Refinement: A Practical Way for Avoiding Overfitting in NAS
- Learning Long-term Visual Dynamics with Region Proposal Interaction Networks
- Training CNNs with Selective Allocation of Channels
- Learning Granularity-Aware Convolutional Neural Network for Fine-Grained Visual Classification
- Layerwise Optimization by Gradient Decomposition for Continual Learning
- Large-Scale Attribute-Object Compositions
- A Simple yet Effective Baseline for Robust Deep Learning with Noisy Labels
- RecSal : Deep Recursive Supervision for Visual Saliency Prediction
- Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition
- Explainable Health Risk Predictor with Transformer-based Medicare Claim Encoder
- DcardNet: Diabetic Retinopathy Classification at Multiple Levels Based on Structural and Angiographic Optical Coherence Tomography
- Effective Data Fusion with Generalized Vegetation Index: Evidence from Land Cover Segmentation in Agriculture
- Accurate Prediction of Free Solvation Energy of Organic Molecules via Graph Attention Network and Message Passing Neural Network from Pairwise Atomistic Interactions
- Exploring Covariate and Concept Shift for Detection and Calibration of Out-of-Distribution Data
- Table-Based Neural Units: Fully Quantizing Networks for Multiply-Free Inference
- Distribution-sensitive Information Retention for Accurate Binary Neural Network
- Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method
- Differentiable Learning-to-Group Channels via Groupable Convolutional Neural Networks
- Greedy Optimization Provably Wins the Lottery: Logarithmic Number of Winning Tickets is Enough
- Exploring Text-transformers in AAAI 2021 Shared Task: COVID-19 Fake News Detection in English
- Doubly Adaptive Scaled Algorithm for Machine Learning Using Second-Order Information
- Analytical Moment Regularizer for Gaussian Robust Networks
- Channel Max Pooling Layer for Fine-Grained Vehicle Classification
- Deep Networks with Fast Retraining
- Continuous-Depth Neural Models for Dynamic Graph Prediction
- Neural network for multi-exponential sound energy decay analysis
- Contextual Recurrent Neural Networks
- Hyperbolic Deep Learning for Chinese Natural Language Understanding
- Optimizing Neural Architecture Search using Limited GPU Time in a Dynamic Search Space: A Gene Expression Programming Approach
- NeRV: Neural Representations for Videos
- Visualizing Adapted Knowledge in Domain Transfer
- Augmented Random Search for Quadcopter Control: An alternative to Reinforcement Learning
- Building a Regular Decision Boundary with Deep Networks
- Annotation-Free Human Sketch Quality Assessment
- MemNet: Memory-Efficiency Guided Neural Architecture Search with Augment-Trim learning
- BNAS v2: Learning Architectures for Binary Networks with Empirical Improvements
- Medical Matting: A New Perspective on Medical Segmentation with Uncertainty
- Long-Short Temporal Contrastive Learning of Video Transformers
- Compressive Visual Representations
- Variance Reduction on General Adaptive Stochastic Mirror Descent
- Reverse engineering learned optimizers reveals known and novel mechanisms
- SMG: A Shuffling Gradient-Based Method with Momentum
- Scaling Law for Recommendation Models: Towards General-purpose User Representations
- All-In-One: Artificial Association Neural Networks
- Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning
- Training Sparse Neural Networks using Compressed Sensing
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- VidFace: A Full-Transformer Solver for Video FaceHallucination with Unaligned Tiny Snapshots
- Deep Learning Methods for Real-time Detection and Analysis of Wagner Ulcer Classification System
- Wide Mean-Field Variational Bayesian Neural Networks Ignore the Data
- Universal Adder Neural Networks
- Embracing the Dark Knowledge: Domain Generalization Using Regularized Knowledge Distillation
- Supervision Accelerates Pre-training in Contrastive Semi-Supervised Learning of Visual Representations
- Learning Hierarchical Graph Neural Networks for Image Clustering
- Exploring the Loss Landscape in Neural Architecture Search
- Aug3D-RPN: Improving Monocular 3D Object Detection by Synthetic Images with Virtual Depth
- Top-1 Solution of Multi-Moments in Time Challenge 2019
- Learning to Structure an Image with Few Colors
- Multi-task problems are not multi-objective
- Wide and Narrow: Video Prediction from Context and Motion
- Arbitrary Marginal Neural Ratio Estimation for Simulation-based Inference
- Balance-Oriented Focal Loss with Linear Scheduling for Anchor Free Object Detection
- Adversarial-Based Knowledge Distillation for Multi-Model Ensemble and Noisy Data Refinement
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- RecNets: Channel-wise Recurrent Convolutional Neural Networks
- Probabilistic Oriented Object Detection in Automotive Radar
- Paying more attention to snapshots of Iterative Pruning: Improving Model Compression via Ensemble Distillation
- Multi-step Estimation for Gradient-based Meta-learning
- Capturing scattered discriminative information using a deep architecture in acoustic scene classification
- Geometric Operator Convolutional Neural Network
- Text-Independent Speaker Verification with Dual Attention Network
- A New Unified Deep Learning Approach with Decomposition-Reconstruction-Ensemble Framework for Time Series Forecasting
- URNet : User-Resizable Residual Networks with Conditional Gating Module
- On the Intrinsic Dimensionality of Image Representations
- H-OWAN: Multi-distorted Image Restoration with Tensor 1x1 Convolution
- RIANN -- A Robust Neural Network Outperforms Attitude Estimation Filters
- Acceleration via Fractal Learning Rate Schedules
- Self-semi-supervised Learning to Learn from NoisyLabeled Data
- GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent
- A Multiclass Boosting Framework for Achieving Fast and Provable Adversarial Robustness
- On the Bias Against Inductive Biases
- Permutation Matters: Anisotropic Convolutional Layer for Learning on Point Clouds
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- Exploring Set Similarity for Dense Self-supervised Representation Learning
- Towards Learning an Unbiased Classifier from Biased Data via Conditional Adversarial Debiasing
- Analysing Dropout and Compounding Errors in Neural Language Models
- Multi-objective Neural Architecture Search with Almost No Training
- Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms
- Supervised Momentum Contrastive Learning for Few-Shot Classification
- Self-Distilled Self-Supervised Representation Learning
- PGT: A Progressive Method for Training Models on Long Videos
- Disturbing Target Values for Neural Network Regularization
- Cross-Channel Intragroup Sparsity Neural Network
- Improving the sample-efficiency of neural architecture search with reinforcement learning
- Toward Runtime-Throttleable Neural Networks
- ExCon: Explanation-driven Supervised Contrastive Learning for Image Classification
- Explore Image Deblurring via Blur Kernel Space
- Understanding the Disharmony between Weight Normalization Family and Weight Decay: shifted Regularizer
- Dual Reconstruction with Densely Connected Residual Network for Single Image Super-Resolution
- AutoShrink: A Topology-aware NAS for Discovering Efficient Neural Architecture
- SpectroscopyNet: Learning to pre-process Spectroscopy Signals without clean data
- Combining learning rate decay and weight decay with complexity gradient descent - Part I
- Trivial or impossible -- dichotomous data difficulty masks model differences (on ImageNet and beyond)
- CondenseNet V2: Sparse Feature Reactivation for Deep Networks
- Unsupervised Embedding Learning from Uncertainty Momentum Modeling
- Histogram of Cell Types: Deep Learning for Automated Bone Marrow Cytology
- Knowledge Transfer Based Fine-grained Visual Classification
- Dense-Resolution Network for Point Cloud Classification and Segmentation
- SIPA: A Simple Framework for Efficient Networks
- GOALS: Gradient-Only Approximations for Line Searches Towards Robust and Consistent Training of Deep Neural Networks
- How Does BN Increase Collapsed Neural Network Filters?
- Pruning Ternary Quantization
- Learning to Profile: User Meta-Profile Network for Few-Shot Learning
- Adaptive Regularization via Residual Smoothing in Deep Learning Optimization
- ReRankMatch: Semi-Supervised Learning with Semantics-Oriented Similarity Representation
- Fully Quantized Image Super-Resolution Networks
- Adversarial Attacks on ML Defense Models Competition
- CRAUM-Net: Contextual Recursive Attention with Uncertainty Modeling for Salient Object Detection
- Multi-dataset Pretraining: A Unified Model for Semantic Segmentation
- Self-supervised Contrastive Learning for EEG-based Sleep Staging
- PipeNet: Selective Modal Pipeline of Fusion Network for Multi-Modal Face Anti-Spoofing
- PHYRE: A New Benchmark for Physical Reasoning
- Re-examination of the Role of Latent Variables in Sequence Modeling
- On SGD's Failure in Practice: Characterizing and Overcoming Stalling
- Convergent Graph Solvers
- Channel DropBlock: An Improved Regularization Method for Fine-Grained Visual Classification
- Rescuing Deep Hashing from Dead Bits Problem
- Hierarchical Auxiliary Learning
- Towards More Efficient and Effective Inference: The Joint Decision of Multi-Participants
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- Network Learning with Local Propagation
- CarneliNet: Neural Mixture Model for Automatic Speech Recognition
- Logit Attenuating Weight Normalization
- Domain Consistency Regularization for Unsupervised Multi-source Domain Adaptive Classification
- Analyze and Design Network Architectures by Recursion Formulas
- Contrastive Representations for Label Noise Require Fine-Tuning
- Bridged Adversarial Training
- On-target Adaptation
- SdcNet: A Computation-Efficient CNN for Object Recognition
- Flexible numerical optimization with ensmallen
- Training Dynamic based data filtering may not work for NLP datasets
- Improving Text-Independent Speaker Verification with Auxiliary Speakers Using Graph
- Distributed Optimization using Heterogeneous Compute Systems
- Contributions to Large Scale Bayesian Inference and Adversarial Machine Learning
- Balanced Masked and Standard Face Recognition
- Efficient Modelling Across Time of Human Actions and Interactions
- Improving Adversarial Robustness for Free with Snapshot Ensemble
- Estimating the Uncertainty of Neural Network Forecasts for Influenza Prevalence Using Web Search Activity
- Building effective deep neural network architectures one feature at a time
- Itsy Bitsy SpiderNet: Fully Connected Residual Network for Fraud Detection
- Rethinking Layer-wise Feature Amounts in Convolutional Neural Network Architectures
- Sinusoidal Flow: A Fast Invertible Autoregressive Flow
- Unlocking the Full Potential of Small Data with Diverse Supervision
- Sub-clusters of Normal Data for Anomaly Detection
- SwGridNet: A Deep Convolutional Neural Network based on Grid Topology for Image Classification
- SelectScale: Mining More Patterns from Images via Selective and Soft Dropout
- Ouroboros: On Accelerating Training of Transformer-Based Language Models
- Neural Architecture Search in Embedding Space
- Network Parameter Learning Using Nonlinear Transforms, Local Representation Goals and Local Propagation Constraints
- Towards Enhancing Fine-grained Details for Image Matting
- Pixel-wise Segmentation of Right Ventricle of Heart
- On estimating gaze by self-attention augmented convolutions
- Zero Training Overhead Portfolios for Learning to Solve Combinatorial Problems
- Vision Pair Learning: An Efficient Training Framework for Image Classification
- Modeling turbulent and self-gravitating fluids with Fourier neural operators
- Searching by Generating: Flexible and Efficient One-Shot NAS with Architecture Generator
- The Best of Both Worlds: a Framework for Combining Degradation Prediction with High Performance Super-Resolution Networks
- Learning Sparse Structured Ensembles with SG-MCMC and Network Pruning
- Generalized Organ Segmentation by Imitating One-shot Reasoning using Anatomical Correlation
- BNAS-v2: Memory-efficient and Performance-collapse-prevented Broad Neural Architecture Search
- Learning Metrics from Mean Teacher: A Supervised Learning Method for Improving the Generalization of Speaker Verification System
- Making Differentiable Architecture Search less local
- Exploration into Translation-Equivariant Image Quantization
- Spatially Attentive Output Layer for Image Classification
- Burst Image Restoration and Enhancement
- Adaptive Attention Span in Computer Vision
- Shape-Oriented Convolution Neural Network for Point Cloud Analysis
- Alternating Synthetic and Real Gradients for Neural Language Modeling
- Improved Robustness of Vision Transformer via PreLayerNorm in Patch Embedding
- MC-SSL0.0: Towards Multi-Concept Self-Supervised Learning
- Low Power In-Memory Implementation of Ternary Neural Networks with Resistive RAM-Based Synapse
- Population Based Training for Data Augmentation and Regularization in Speech Recognition
- Modulated binary cliquenet
- Be Your Own Best Competitor! Multi-Branched Adversarial Knowledge Transfer
- Beyond Attributes: Adversarial Erasing Embedding Network for Zero-shot Learning
- Analysis of Atomistic Representations Using Weighted Skip-Connections
- MMCGAN: Generative Adversarial Network with Explicit Manifold Prior
- Sequential Feature Filtering Classifier
- AntiDote: Attention-based Dynamic Optimization for Neural Network Runtime Efficiency
- Asymmetric Variational Autoencoders
- Privacy Preserving Recalibration under Domain Shift
- AAG: Self-Supervised Representation Learning by Auxiliary Augmentation with GNT-Xent Loss
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning