Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
arXiv:1706.02677
Abstract
Deep learning thrives with large neural networks and large datasets. However, larger networks and larger datasets result in longer training times that impede research and development progress. Distributed synchronous SGD offers a potential solution to this problem by dividing SGD minibatches over a pool of parallel workers. Yet to make this scheme efficient, the per-worker workload must be large, which implies nontrivial growth in the SGD minibatch size. In this paper, we empirically show that on the ImageNet dataset large minibatches cause optimization difficulties, but when these are addressed the trained networks exhibit good generalization. Specifically, we show no loss of accuracy when training with large minibatch sizes up to 8192 images. To achieve this result, we adopt a hyper-parameter-free linear scaling rule for adjusting learning rates as a function of minibatch size and develop a new warmup scheme that overcomes optimization challenges early in training. With these simple techniques, our Caffe2-based system trains ResNet-50 with a minibatch size of 8192 on 256 GPUs in one hour, while matching small minibatch accuracy. Using commodity hardware, our implementation achieves ~90% scaling efficiency when moving from 8 to 256 GPUs. Our findings enable training visual recognition models on internet-scale data with high efficiency.
Tech report (v2: correct typos)
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Natural Language Processing (almost) from Scratch
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- One weird trick for parallelizing convolutional neural networks
- Revisiting Distributed Synchronous SGD
Cited by in corpus (934)
- A Simple Framework for Contrastive Learning of Visual Representations
- Bootstrap your own latent: A new approach to self-supervised Learning
- YOLOX: Exceeding YOLO Series in 2021
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Unsupervised Representation Learning by Predicting Image Rotations
- Is Space-Time Attention All You Need for Video Understanding?
- Vision Transformers for Single Image Dehazing
- Dota 2 with Large Scale Deep Reinforcement Learning
- Momentum Contrast for Unsupervised Visual Representation Learning
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- Towards Federated Learning at Scale: System Design
- A Survey on Distributed Machine Learning
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- MMDetection: Open MMLab Detection Toolbox and Benchmark
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification
- SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
- On the Variance of the Adaptive Learning Rate and Beyond
- Visualizing the Loss Landscape of Neural Nets
- Enhancing Graph Neural Network-based Fraud Detectors against Camouflaged Fraudsters
- Tune: A Research Platform for Distributed Model Selection and Training
- DeePMD-kit v2: A software package for Deep Potential models
- Don't Decay the Learning Rate, Increase the Batch Size
- AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
- Horovod: fast and easy distributed deep learning in TensorFlow
- Large Batch Training of Convolutional Networks
- Regularizing and Optimizing LSTM Language Models
- Exploring Simple Siamese Representation Learning
- Train longer, generalize better: closing the generalization gap in large batch training of neural networks
- Precision Health Data: Requirements, Challenges and Existing Techniques for Data Security and Privacy
- Lookahead Optimizer: k steps forward, 1 step back
- Sustainable AI: Environmental Implications, Challenges and Opportunities
- Early Convolutions Help Transformers See Better
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Revisiting Small Batch Training for Deep Neural Networks
- Centralized Feature Pyramid for Object Detection
- Rosetta: Large scale system for text detection and recognition in images
- Billion-scale semi-supervised learning for image classification
- Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
- Tent: Fully Test-time Adaptation by Entropy Minimization
- Deep Neural Networks to Enable Real-time Multimessenger Astrophysics
- Training Tips for the Transformer Model
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted Instances
- DARTS+: Improved Differentiable Architecture Search with Early Stopping
- High-Performance Large-Scale Image Recognition Without Normalization
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning
- Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
- Detecting Anemia from Retinal Fundus Images
- Federated Variance-Reduced Stochastic Gradient Descent with Robustness to Byzantine Attacks
- Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss
- Deep Residual Learning in Spiking Neural Networks
- Object-Centric Learning with Slot Attention
- A Closer Look at Spatiotemporal Convolutions for Action Recognition
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
- Don't Use Large Mini-Batches, Use Local SGD
- Revisiting ResNets: Improved Training and Scaling Strategies
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Local SGD Converges Fast and Communicates Little
- Masked Autoencoders Are Scalable Vision Learners
- MHSA-Net: Multi-Head Self-Attention Network for Occluded Person Re-Identification
- ShakeDrop Regularization for Deep Residual Learning
- On Empirical Comparisons of Optimizers for Deep Learning
- Natural Adversarial Examples
- Neural Architecture Transfer
- MLPerf Training Benchmark
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
- Contrastive Masked Autoencoders are Stronger Vision Learners
- Bringing AI To Edge: From Deep Learning's Perspective
- A Field Guide to Federated Optimization
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- Adafactor: Adaptive Learning Rates with Sublinear Memory Cost
- Differentiable Learning-to-Normalize via Switchable Normalization
- Pythia v0.1: the Winning Entry to the VQA Challenge 2018
- Audiovisual SlowFast Networks for Video Recognition
- Measuring the Effects of Data Parallelism on Neural Network Training
- Large scale distributed neural network training through online distillation
- Non-local Neural Networks
- Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution
- Database Meets Deep Learning: Challenges and Opportunities
- ReZero is All You Need: Fast Convergence at Large Depth
- Bag of Freebies for Training Object Detection Neural Networks
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Neonatal seizure detection from raw multi-channel EEG using a fully convolutional architecture
- PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization
- Beyond Data and Model Parallelism for Deep Neural Networks
- Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data
- AutoSlim: Towards One-Shot Architecture Search for Channel Numbers
- SlowFast Networks for Video Recognition
- Bag of Tricks for Image Classification with Convolutional Neural Networks
- BigDL: A Distributed Deep Learning Framework for Big Data
- Uncovering the Limits of Adversarial Training against Norm-Bounded Adversarial Examples
- Stabilizing the Lottery Ticket Hypothesis
- Classification Accuracy Score for Conditional Generative Models
- Pyramidal Convolution: Rethinking Convolutional Neural Networks for Visual Recognition
- Self-supervised Pretraining of Visual Features in the Wild
- Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey
- Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation
- An Analysis of Neural Language Modeling at Multiple Scales
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- Optimization for deep learning: theory and algorithms
- Compact Generalized Non-local Network
- Class-Balanced Loss Based on Effective Number of Samples
- Deep Learning in Mobile and Wireless Networking: A Survey
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Adaptive Federated Optimization
- Practical Deep Learning with Bayesian Principles
- Towards Explaining the Regularization Effect of Initial Large Learning Rate in Training Neural Networks
- RetinaMask: Learning to predict masks improves state-of-the-art single-shot detection for free
- Scaling Distributed Machine Learning with In-Network Aggregation
- Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
- A Comprehensive Study of Deep Video Action Recognition
- Recent Advances in Object Detection in the Age of Deep Convolutional Neural Networks
- Fixup Initialization: Residual Learning Without Normalization
- Review: Deep Learning in Electron Microscopy
- Fixing Data Augmentation to Improve Adversarial Robustness
- AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
- Linear Mode Connectivity and the Lottery Ticket Hypothesis
- The Conditional Entropy Bottleneck
- Rethinking ImageNet Pre-training
- Sewer-ML: A Multi-Label Sewer Defect Classification Dataset and Benchmark
- Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training
- How to train your neural ODE: the world of Jacobian and kinetic regularization
- PrivFT: Private and Fast Text Classification with Homomorphic Encryption
- Training Quantized Nets: A Deeper Understanding
- High-Fidelity Image Generation With Fewer Labels
- Accelerated Methods for Deep Reinforcement Learning
- Automatic Perturbation Analysis for Scalable Certified Robustness and Beyond
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster Inference
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
- Incomplete Descriptor Mining with Elastic Loss for Person Re-Identification
- Selection via Proxy: Efficient Data Selection for Deep Learning
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- The Non-IID Data Quagmire of Decentralized Machine Learning
- A novel Region of Interest Extraction Layer for Instance Segmentation
- Self-Adaptive Training: beyond Empirical Risk Minimization
- AP-Loss for Accurate One-Stage Object Detection
- Partial Connection Based on Channel Attention for Differentiable Neural Architecture Search
- Dynamic Model Pruning with Feedback
- Quasi-hyperbolic momentum and Adam for deep learning
- Video Classification with Channel-Separated Convolutional Networks
- Cluster-Level Contrastive Learning for Emotion Recognition in Conversations
- Adversarial Self-Supervised Contrastive Learning
- Feature Denoising for Improving Adversarial Robustness
- Efficient Self-supervised Vision Transformers for Representation Learning
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- Exploring Randomly Wired Neural Networks for Image Recognition
- Spatiotemporal Contrastive Video Representation Learning
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Labelling unlabelled videos from scratch with multi-modal self-supervision
- Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
- Deep Polynomial Neural Networks
- Tencent ML-Images: A Large-Scale Multi-Label Image Database for Visual Representation Learning
- BBN: Bilateral-Branch Network with Cumulative Learning for Long-Tailed Visual Recognition
- Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
- GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent
- Escaping Saddles with Stochastic Gradients
- Testing Robustness Against Unforeseen Adversaries
- A Mask Attention Interaction and Scale Enhancement Network for SAR Ship Instance Segmentation
- Benchmarking Detection Transfer Learning with Vision Transformers
- Analyzing Human-Human Interactions: A Survey
- Fast Federated Learning by Balancing Communication Trade-Offs
- Lightweight Multi-Branch Network for Person Re-Identification
- Massively Distributed SGD: ImageNet/ResNet-50 Training in a Flash
- Yet Another Accelerated SGD: ResNet-50 Training on ImageNet in 74.7 seconds
- Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
- Revisiting Self-Supervised Visual Representation Learning
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
- A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
- Communication optimization strategies for distributed deep neural network training: A survey
- Natural Compression for Distributed Deep Learning
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss
- A System for Massively Parallel Hyperparameter Tuning
- A Simple Proximal Stochastic Gradient Method for Nonsmooth Nonconvex Optimization
- Asynchronous Decentralized Parallel Stochastic Gradient Descent
- Instance adaptive adversarial training: Improved accuracy tradeoffs in neural nets
- Differentially Private Learning Needs Better Features (or Much More Data)
- FetchSGD: Communication-Efficient Federated Learning with Sketching
- Augmentation for small object detection
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- Rethinking the Hyperparameters for Fine-tuning
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
- FishNet: A Versatile Backbone for Image, Region, and Pixel Level Prediction
- ChainerMN: Scalable Distributed Deep Learning Framework
- A Progressive Batching L-BFGS Method for Machine Learning
- Near-Optimal Sparse Allreduce for Distributed Deep Learning
- Multiscale Vision Transformers
- Relation-Aware Graph Attention Network for Visual Question Answering
- Training Neural Response Selection for Task-Oriented Dialogue Systems
- WNGrad: Learn the Learning Rate in Gradient Descent
- SELF: Learning to Filter Noisy Labels with Self-Ensembling
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- Evolving Normalization-Activation Layers
- i-Mix: A Domain-Agnostic Strategy for Contrastive Representation Learning
- Aligning Pretraining for Detection via Object-Level Contrastive Learning
- Rotate to Attend: Convolutional Triplet Attention Module
- PruneTrain: Fast Neural Network Training by Dynamic Sparse Model Reconfiguration
- LIT-Former: Linking In-plane and Through-plane Transformers for Simultaneous CT Image Denoising and Deblurring
- Long-tail learning via logit adjustment
- Variance-based Gradient Compression for Efficient Distributed Deep Learning
- Restructuring Batch Normalization to Accelerate CNN Training
- Insensitive Stochastic Gradient Twin Support Vector Machine for Large Scale Problems
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
- Unsupervised Deep Representation Learning and Few-Shot Classification of PolSAR Images
- Federated Learning with Unbiased Gradient Aggregation and Controllable Meta Updating
- Augment your batch: better training with larger batches
- An Improved Baseline for Sentence-level Relation Extraction
- MEAL V2: Boosting Vanilla ResNet-50 to 80%+ Top-1 Accuracy on ImageNet without Tricks
- Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks
- Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments
- EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa
- Scalable Distributed DNN Training using TensorFlow and CUDA-Aware MPI: Characterization, Designs, and Performance Evaluation
- Drawing Early-Bird Tickets: Towards More Efficient Training of Deep Networks
- Optimizing Network Performance for Distributed DNN Training on GPU Clusters: ImageNet/AlexNet Training in 1.5 Minutes
- Straggler-aware Distributed Learning: Communication Computation Latency Trade-off
- Model Rubik's Cube: Twisting Resolution, Depth and Width for TinyNets
- Deep Learning at Scale for the Construction of Galaxy Catalogs in the Dark Energy Survey
- Accelerating SGD with momentum for over-parameterized learning
- Machine learning astrophysics from 21 cm lightcones: impact of network architectures and signal contamination
- ImageNet Training in Minutes
- Decoupled Classification Refinement: Hard False Positive Suppression for Object Detection
- Four Things Everyone Should Know to Improve Batch Normalization
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- Bounding Box Regression with Uncertainty for Accurate Object Detection
- Traditional and Heavy-Tailed Self Regularization in Neural Network Models
- On the Origin of Implicit Regularization in Stochastic Gradient Descent
- Understanding Batch Normalization
- TinyissimoYOLO: A Quantized, Low-Memory Footprint, TinyML Object Detection Network for Low Power Microcontrollers
- RedSync : Reducing Synchronization Traffic for Distributed Deep Learning
- FcaNet: Frequency Channel Attention Networks
- Cream of the Crop: Distilling Prioritized Paths For One-Shot Neural Architecture Search
- BigNAS: Scaling Up Neural Architecture Search with Big Single-Stage Models
- TEA: Temporal Excitation and Aggregation for Action Recognition
- Semi-Dynamic Load Balancing: Efficient Distributed Learning in Non-Dedicated Environments
- Federated Learning with Buffered Asynchronous Aggregation
- DD-PPO: Learning Near-Perfect PointGoal Navigators from 2.5 Billion Frames
- Knowledge Distillation via Route Constrained Optimization
- Massive MIMO Channel Prediction Via Meta-Learning and Deep Denoising: Is a Small Dataset Enough?
- Compounding the Performance Improvements of Assembled Techniques in a Convolutional Neural Network
- Distribution-Balanced Loss for Multi-Label Classification in Long-Tailed Datasets
- TensorMask: A Foundation for Dense Object Segmentation
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- Scale MLPerf-0.6 models on Google TPU-v3 Pods
- Deformable Kernels: Adapting Effective Receptive Fields for Object Deformation
- SGD: General Analysis and Improved Rates
- Online Normalization for Training Neural Networks
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- OneFlow: Redesign the Distributed Deep Learning Framework from Scratch
- MultiGrain: a unified image embedding for classes and instances
- Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
- Large batch size training of neural networks with adversarial training and second-order information
- Long-Term Feature Banks for Detailed Video Understanding
- The Practicality of Stochastic Optimization in Imaging Inverse Problems
- Exploring Self-attention for Image Recognition
- Radioactive data: tracing through training
- Towards Efficient Training for Neural Network Quantization
- TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
- Towards Better Accuracy-efficiency Trade-offs: Divide and Co-training
- VarifocalNet: An IoU-aware Dense Object Detector
- Declarative Recursive Computation on an RDBMS, or, Why You Should Use a Database For Distributed Machine Learning
- Linearly Converging Error Compensated SGD
- torchgpipe: On-the-fly Pipeline Parallelism for Training Giant Models
- Distributed Training with Heterogeneous Data: Bridging Median- and Mean-Based Algorithms
- Small-GAN: Speeding Up GAN Training Using Core-sets
- Moniqua: Modulo Quantized Communication in Decentralized SGD
- Improving accuracy and speeding up Document Image Classification through parallel systems
- Tesseract: Parallelize the Tensor Parallelism Efficiently
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- Stochastic, Distributed and Federated Optimization for Machine Learning
- NVIDIA SimNet^{TM}: an AI-accelerated multi-physics simulation framework
- Single-stream CNN with Learnable Architecture for Multi-source Remote Sensing Data
- An Efficient Statistical-based Gradient Compression Technique for Distributed Training Systems
- Hierarchical Federated Learning through LAN-WAN Orchestration
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- Decentralized Deep Learning with Arbitrary Communication Compression
- The Difficulty of Training Sparse Neural Networks
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints
- Lost in Pruning: The Effects of Pruning Neural Networks beyond Test Accuracy
- The Impact of the Mini-batch Size on the Variance of Gradients in Stochastic Gradient Descent
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- MST: Masked Self-Supervised Transformer for Visual Representation
- SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning
- Scaling Wide Residual Networks for Panoptic Segmentation
- MegDet: A Large Mini-Batch Object Detector
- Chainer: A Deep Learning Framework for Accelerating the Research Cycle
- EvoPose2D: Pushing the Boundaries of 2D Human Pose Estimation using Accelerated Neuroevolution with Weight Transfer
- Light-Weight RetinaNet for Object Detection
- Understanding the Difficulty of Training Transformers
- On Feature Normalization and Data Augmentation
- Multi-Task and Multi-Modal Learning for RGB Dynamic Gesture Recognition
- Vehicle Attribute Recognition by Appearance: Computer Vision Methods for Vehicle Type, Make and Model Classification
- Communication-Efficient Edge AI: Algorithms and Systems
- SA-Net: Shuffle Attention for Deep Convolutional Neural Networks
- Rethinking "Batch" in BatchNorm
- Is Label Smoothing Truly Incompatible with Knowledge Distillation: An Empirical Study
- Towards Scalable Distributed Training of Deep Learning on Public Cloud Clusters
- Understanding Generalization through Visualizations
- Transfer of Adversarial Robustness Between Perturbation Types
- CARS: Continuous Evolution for Efficient Neural Architecture Search
- Automated scoring of pre-REM sleep in mice with deep learning
- Improved Residual Networks for Image and Video Recognition
- Deep Leakage from Gradients
- CT-Net: Channel Tensorization Network for Video Classification
- Direction Concentration Learning: Enhancing Congruency in Machine Learning
- Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling
- Equalization Loss for Long-Tailed Object Recognition
- AdaX: Adaptive Gradient Descent with Exponential Long Term Memory
- FEDZIP: A Compression Framework for Communication-Efficient Federated Learning
- Data Movement Is All You Need: A Case Study on Optimizing Transformers
- WebFace260M: A Benchmark Unveiling the Power of Million-Scale Deep Face Recognition
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
- Instance Shadow Detection with A Single-Stage Detector
- Masked Face Recognition Challenge: The WebFace260M Track Report
- Learning from History for Byzantine Robust Optimization
- Accurate Face Detection for High Performance
- Deep neural networks-based denoising models for CT imaging and their efficacy
- Agriculture-Vision: A Large Aerial Image Database for Agricultural Pattern Analysis
- Consensus Control for Decentralized Deep Learning
- Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron
- Batch DropBlock Network for Person Re-identification and Beyond
- Temporal Pyramid Network for Action Recognition
- Disentangling Label Distribution for Long-tailed Visual Recognition
- Learning Imbalanced Datasets with Maximum Margin Loss
- An ensemble-based approach by fine-tuning the deep transfer learning models to classify pneumonia from chest X-ray images
- NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
- Recent advances in deep learning theory
- Disentangled Variational Autoencoder for Emotion Recognition in Conversations
- DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
- Self-Supervised Learning with Kernel Dependence Maximization
- An Effective Anti-Aliasing Approach for Residual Networks
- Margin Matters: Towards More Discriminative Deep Neural Network Embeddings for Speaker Recognition
- Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image Retrieval
- SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization
- Model Slicing for Supporting Complex Analytics with Elastic Inference Cost and Resource Constraints
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
- TF-Replicator: Distributed Machine Learning for Researchers
- Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
- Gradient Diversity: a Key Ingredient for Scalable Distributed Learning
- Reinforced Neighborhood Selection Guided Multi-Relational Graph Neural Networks
- Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
- Maximizing Parallelism in Distributed Training for Huge Neural Networks
- On the Outsized Importance of Learning Rates in Local Update Methods
- BYOL works even without batch statistics
- Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
- Parle: parallelizing stochastic gradient descent
- On Scale-out Deep Learning Training for Cloud and HPC
- On the Utility of Gradient Compression in Distributed Training Systems
- Ensemble of ACCDOA- and EINV2-based Systems with D3Nets and Impulse Response Simulation for Sound Event Localization and Detection
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
- A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Augmented Skeleton Based Contrastive Action Learning with Momentum LSTM for Unsupervised Action Recognition
- Connecting optical morphology, environment, and HI mass fraction for low-redshift galaxies using deep learning
- Communication trade-offs for synchronized distributed SGD with large step size
- Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- Dynamic Parameter Allocation in Parameter Servers
- Shape Matters: Understanding the Implicit Bias of the Noise Covariance
- M2m: Imbalanced Classification via Major-to-minor Translation
- Semantic Redundancies in Image-Classification Datasets: The 10% You Don't Need
- Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
- Towards Crowdsourced Training of Large Neural Networks using Decentralized Mixture-of-Experts
- AdaBits: Neural Network Quantization with Adaptive Bit-Widths
- MetaSAug: Meta Semantic Augmentation for Long-Tailed Visual Recognition
- Performance Modeling and Evaluation of Distributed Deep Learning Frameworks on GPUs
- CalibrationPhys: Self-supervised Video-based Heart and Respiratory Rate Measurements by Calibrating Between Multiple Cameras
- Graph-Based Global Reasoning Networks
- of two-dimensional electron gas: a neural canonical transformation study
- GTA: Global Temporal Attention for Video Action Understanding
- Communication-efficient distributed SGD with Sketching
- Robust Learning Under Label Noise With Iterative Noise-Filtering
- GeoCLR: Georeference Contrastive Learning for Efficient Seafloor Image Interpretation
- Improving Semi-supervised Federated Learning by Reducing the Gradient Diversity of Models
- Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks
- Revisiting Knowledge Distillation via Label Smoothing Regularization
- Estimation of discrete choice models with hybrid stochastic adaptive batch size algorithms
- Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
- Sparse Communication for Training Deep Networks
- Influence Functions in Deep Learning Are Fragile
- PFDet: 2nd Place Solution to Open Images Challenge 2018 Object Detection Track
- Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
- Large Scale Language Modeling: Converging on 40GB of Text in Four Hours
- Modality-Pairing Learning for Brain Tumor Segmentation
- CLAR: Contrastive Learning of Auditory Representations
- MARINA: Faster Non-Convex Distributed Learning with Compression
- Pre-Trained Models: Past, Present and Future
- Pruning-aware Sparse Regularization for Network Pruning
- NODE-GAM: Neural Generalized Additive Model for Interpretable Deep Learning
- The Architectural Implications of Facebook's DNN-based Personalized Recommendation
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Feature Intertwiner for Object Detection
- IoU-aware Single-stage Object Detector for Accurate Localization
- Puzzle-AE: Novelty Detection in Images through Solving Puzzles
- Rethinking Channel Dimensions for Efficient Model Design
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
- Layer-wise Adaptive Gradient Sparsification for Distributed Deep Learning with Convergence Guarantees
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources
- NetReduce: RDMA-Compatible In-Network Reduction for Distributed DNN Training Acceleration
- DurIAN-SC: Duration Informed Attention Network based Singing Voice Conversion System
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
- The Convergence of Sparsified Gradient Methods
- On the Ineffectiveness of Variance Reduced Optimization for Deep Learning
- Bias-Variance Reduced Local SGD for Less Heterogeneous Federated Learning
- Taming Momentum in a Distributed Asynchronous Environment
- KHAN: Knowledge-Aware Hierarchical Attention Networks for Accurate Political Stance Prediction
- HPC AI500: The Methodology, Tools, Roofline Performance Models, and Metrics for Benchmarking HPC AI Systems
- Pruning via Iterative Ranking of Sensitivity Statistics
- Group Ensemble: Learning an Ensemble of ConvNets in a single ConvNet
- Stochastic Nonconvex Optimization with Large Minibatches
- To Talk or to Work: Flexible Communication Compression for Energy Efficient Federated Learning over Heterogeneous Mobile Edge Devices
- Bag of Tricks for Adversarial Training
- Adam: A Stochastic Method with Adaptive Variance Reduction
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- Boundary-preserving Mask R-CNN
- Neural canonical transformations for vibrational spectra of molecules
- Domain-specific Communication Optimization for Distributed DNN Training
- Hierarchical Contrastive Motion Learning for Video Action Recognition
- Spectral Feature Transformation for Person Re-identification
- Narrowing the Gap: Improved Detector Training with Noisy Location Annotations
- BlueFog: Make Decentralized Algorithms Practical for Optimization and Deep Learning
- Boosting Distributed Machine Learning Training Through Loss-tolerant Transmission Protocol
- Hierarchical Opacity Propagation for Image Matting
- Iterative Normalization: Beyond Standardization towards Efficient Whitening
- p-Meta: Towards On-device Deep Model Adaptation
- Bayesian Deep Learning via Subnetwork Inference
- Communication-Efficient Decentralized Learning with Sparsification and Adaptive Peer Selection
- Can weight sharing outperform random architecture search? An investigation with TuNAS
- Doubly Contrastive Deep Clustering
- Momentum^2 Teacher: Momentum Teacher with Momentum Statistics for Self-Supervised Learning
- ACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-Shot
- Scalable Training of Trustworthy and Energy-Efficient Predictive Graph Foundation Models for Atomistic Materials Modeling: A Case Study with HydraGNN
- Learning the Relation between Similarity Loss and Clustering Loss in Self-Supervised Learning
- Farewell to Mutual Information: Variational Distillation for Cross-Modal Person Re-Identification
- Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization
- Extended Batch Normalization
- Neural Network Pruning with Residual-Connections and Limited-Data
- Training-Time-Friendly Network for Real-Time Object Detection
- An Empirical Study of Large-Batch Stochastic Gradient Descent with Structured Covariance Noise
- HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
- SparCML: High-Performance Sparse Communication for Machine Learning
- Accelerating DNN Training in Wireless Federated Edge Learning Systems
- GradiVeQ: Vector Quantization for Bandwidth-Efficient Gradient Aggregation in Distributed CNN Training
- On the use of neural networks for the structural characterization of polymeric porous materials
- TF-SepNet: An Efficient 1D Kernel Design in CNNs for Low-Complexity Acoustic Scene Classification
- Towards Foundation Models for Materials Science: The Open MatSci ML Toolkit
- LiftFormer: 3D Human Pose Estimation using attention models
- A3D: Adaptive 3D Networks for Video Action Recognition
- Accelerated Large Batch Optimization of BERT Pretraining in 54 minutes
- Block-diagonal Hessian-free Optimization for Training Neural Networks
- Straggler-Resilient Distributed Machine Learning with Dynamic Backup Workers
- On Interaction Between Augmentations and Corruptions in Natural Corruption Robustness
- EcoNAS: Finding Proxies for Economical Neural Architecture Search
- Deep Affinity Net: Instance Segmentation via Affinity
- Mirror Descent View for Neural Network Quantization
- Exploring the limits of Concurrency in ML Training on Google TPUs
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
- Instance Segmentation of Visible and Occluded Regions for Finding and Picking Target from a Pile of Objects
- Spatially Consistent Representation Learning
- Learning Singing From Speech
- Training Sound Event Detection On A Heterogeneous Dataset
- Improved Analysis of Clipping Algorithms for Non-convex Optimization
- Batch Group Normalization
- Cyclic Differentiable Architecture Search
- Rethinking Dilated Convolution for Real-time Semantic Segmentation
- The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism
- IROF: a low resource evaluation metric for explanation methods
- DR Loss: Improving Object Detection by Distributional Ranking
- Parallel Training of Deep Networks with Local Updates
- A vector quantized masked autoencoder for audiovisual speech emotion recognition
- Revisiting the Sibling Head in Object Detector
- Simpler, Faster, Stronger: Breaking The log-K Curse On Contrastive Learners With FlatNCE
- Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Partial Order Pruning: for Best Speed/Accuracy Trade-off in Neural Architecture Search
- Learning Rate Annealing Can Provably Help Generalization, Even for Convex Problems
- Learning from Temporal Gradient for Semi-supervised Action Recognition
- ScaDLES: Scalable Deep Learning over Streaming data at the Edge
- IoU-uniform R-CNN: Breaking Through the Limitations of RPN
- A Distributed Synchronous SGD Algorithm with Global Top- Sparsification for Low Bandwidth Networks
- Per-Tensor Fixed-Point Quantization of the Back-Propagation Algorithm
- Face Detection with Feature Pyramids and Landmarks
- end-to-end training of a large vocabulary end-to-end speech recognition system
- Multi-task learning for electronic structure to predict and explore molecular potential energy surfaces
- Automated Classification of Sleep Stages and EEG Artifacts in Mice with Deep Learning
- Distributed Stochastic Algorithms for High-rate Streaming Principal Component Analysis
- Online Hyper-parameter Learning for Auto-Augmentation Strategy
- Interpretable agent communication from scratch (with a generic visual processor emerging on the side)
- MixSiam: A Mixture-based Approach to Self-supervised Representation Learning
- Auxo: Efficient Federated Learning via Scalable Client Clustering
- Taxonomizing local versus global structure in neural network loss landscapes
- Understanding Training Efficiency of Deep Learning Recommendation Models at Scale
- FedPAGE: A Fast Local Stochastic Gradient Method for Communication-Efficient Federated Learning
- Broadcasted Residual Learning for Efficient Keyword Spotting
- A(DP)SGD: Asynchronous Decentralized Parallel Stochastic Gradient Descent with Differential Privacy
- Multi-Scale Feature Aggregation by Cross-Scale Pixel-to-Region Relation Operation for Semantic Segmentation
- A DAG Model of Synchronous Stochastic Gradient Descent in Distributed Deep Learning
- Disentangling Adaptive Gradient Methods from Learning Rates
- Scaling Ensemble Distribution Distillation to Many Classes with Proxy Targets
- Is SGD a Bayesian sampler? Well, almost
- TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Layered SGD: A Decentralized and Synchronous SGD Algorithm for Scalable Deep Neural Network Training
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
- Heterogeneity-Aware Asynchronous Decentralized Training
- Accelerating CNN Training by Pruning Activation Gradients
- SSN: Learning Sparse Switchable Normalization via SparsestMax
- BaPipe: Exploration of Balanced Pipeline Parallelism for DNN Training
- InfoCNF: An Efficient Conditional Continuous Normalizing Flow with Adaptive Solvers
- Automated Learning Rate Scheduler for Large-batch Training
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- Study on the Large Batch Size Training of Neural Networks Based on the Second Order Gradient
- BroadFace: Looking at Tens of Thousands of People at Once for Face Recognition
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- Manipulating SGD with Data Ordering Attacks
- Powerpropagation: A sparsity inducing weight reparameterisation
- DIFER: Differentiable Automated Feature Engineering
- Robust Training in High Dimensions via Block Coordinate Geometric Median Descent
- Multi-node Bert-pretraining: Cost-efficient Approach
- A Proof of Useful Work for Artificial Intelligence on the Blockchain
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- Proxy-Normalizing Activations to Match Batch Normalization while Removing Batch Dependence
- Self-supervised Motion Learning from Static Images
- Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent
- Regularizing Neural Networks via Adversarial Model Perturbation
- Training a Fully Convolutional Neural Network to Route Integrated Circuits
- Unsupervised Object-Level Representation Learning from Scene Images
- Label Noise SGD Provably Prefers Flat Global Minimizers
- Beyond Human-Level Accuracy: Computational Challenges in Deep Learning
- 1st Place Solutions of Waymo Open Dataset Challenge 2020 -- 2D Object Detection Track
- Real-world Mapping of Gaze Fixations Using Instance Segmentation for Road Construction Safety Applications
- Domain Generalization on Efficient Acoustic Scene Classification using Residual Normalization
- Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling
- You Only Look One-level Feature
- AI Enabling Technologies: A Survey
- Do Normalization Layers in a Deep ConvNet Really Need to Be Distinct?
- A Loss Curvature Perspective on Training Instability in Deep Learning
- A Stochastic Extra-Step Quasi-Newton Method for Nonsmooth Nonconvex Optimization
- Revitalizing CNN Attentions via Transformers in Self-Supervised Visual Representation Learning
- Rethinking Curriculum Learning with Incremental Labels and Adaptive Compensation
- Spherical Motion Dynamics: Learning Dynamics of Neural Network with Normalization, Weight Decay, and SGD
- AOGNets: Compositional Grammatical Architectures for Deep Learning
- Neural networks for semantic segmentation of historical city maps: Cross-cultural performance and the impact of figurative diversity
- Stochastic Training is Not Necessary for Generalization
- Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
- Instance Shadow Detection
- Dynamic learning rate using Mutual Information
- Rethinking Training from Scratch for Object Detection
- Self-Adaptive Training: Bridging Supervised and Self-Supervised Learning
- Representation Learning via Consistent Assignment of Views to Clusters
- HierTrain: Fast Hierarchical Edge AI Learning with Hybrid Parallelism in Mobile-Edge-Cloud Computing
- Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
- Natural Image Matting via Guided Contextual Attention
- AgEBO-Tabular: Joint Neural Architecture and Hyperparameter Search with Autotuned Data-Parallel Training for Tabular Data
- Asynchronous Interaction Aggregation for Action Detection
- REX: Revisiting Budgeted Training with an Improved Schedule
- Pufferfish: Communication-efficient Models At No Extra Cost
- SimTriplet: Simple Triplet Representation Learning with a Single GPU
- S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima
- MUXConv: Information Multiplexing in Convolutional Neural Networks
- Points as Queries: Weakly Semi-supervised Object Detection by Points
- A Scaling Law for Synthetic-to-Real Transfer: How Much Is Your Pre-training Effective?
- Distilling Virtual Examples for Long-tailed Recognition
- Training Deep Neural Networks Without Batch Normalization
- Fire SSD: Wide Fire Modules based Single Shot Detector on Edge Device
- Curved Text Detection in Natural Scene Images with Semi- and Weakly-Supervised Learning
- Generalized Zero-Shot Domain Adaptation via Coupled Conditional Variational Autoencoders
- Enhancing Transformers without Self-supervised Learning: A Loss Landscape Perspective in Sequential Recommendation
- SGNet: A Super-class Guided Network for Image Classification and Object Detection
- Pseudo-IoU: Improving Label Assignment in Anchor-Free Object Detection
- Unsupervised Deep Feature Transfer for Low Resolution Image Classification
- HyPar-Flow: Exploiting MPI and Keras for Scalable Hybrid-Parallel DNN Training using TensorFlow
- Democratizing Production-Scale Distributed Deep Learning
- Implicit Regularization and Convergence for Weight Normalization
- Variance reduction for Riemannian non-convex optimization with batch size adaptation
- Parallel Complexity of Forward and Backward Propagation
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
- Compressing Gradient Optimizers via Count-Sketches
- A Multigrid Method for Efficiently Training Video Models
- Deep MRI Reconstruction with Radial Subsampling
- iffDetector: Inference-aware Feature Filtering for Object Detection
- The Impact of GPU DVFS on the Energy and Performance of Deep Learning: an Empirical Study
- Multi-Task Self-Training for Learning General Representations
- Novel Human-Object Interaction Detection via Adversarial Domain Generalization
- Sketchy Empirical Natural Gradient Methods for Deep Learning
- Robust and On-the-fly Dataset Denoising for Image Classification
- Newton-ADMM: A Distributed GPU-Accelerated Optimizer for Multiclass Classification Problems
- Video-based Person Re-identification via 3D Convolutional Networks and Non-local Attention
- Self-supervised Learning of Rotation-invariant 3D Point Set Features using Transformer and its Self-distillation
- Towards Understanding Iterative Magnitude Pruning: Why Lottery Tickets Win
- Learning by Turning: Neural Architecture Aware Optimisation
- ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training
- SeesawFaceNets: sparse and robust face verification model for mobile platform
- Multi-scale Aggregation R-CNN for 2D Multi-person Pose Estimation
- MoGA: Searching Beyond MobileNetV3
- Concatenated Feature Pyramid Network for Instance Segmentation
- Speeding up Deep Learning with Transient Servers
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise Decomposition
- Stanza: Layer Separation for Distributed Training in Deep Learning
- Large-Scale Attribute-Object Compositions
- Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error
- On Feature Decorrelation in Self-Supervised Learning
- PSGAN: A Generative Adversarial Network for Remote Sensing Image Pan-Sharpening
- Gradient Energy Matching for Distributed Asynchronous Gradient Descent
- Learnable Companding Quantization for Accurate Low-bit Neural Networks
- Stagewise Enlargement of Batch Size for SGD-based Learning
- Better SGD using Second-order Momentum
- ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis
- Distributed Deep Learning in Open Collaborations
- SWAD: Domain Generalization by Seeking Flat Minima
- EqCo: Equivalent Rules for Self-supervised Contrastive Learning
- Dynamic Sparse Graph for Efficient Deep Learning
- Learning with Hierarchical Complement Objective
- Low-Rank Training of Deep Neural Networks for Emerging Memory Technology
- VirtualFlow: Decoupling Deep Learning Models from the Underlying Hardware
- Optimal Mini-Batch Size Selection for Fast Gradient Descent
- Online Continual Learning with Natural Distribution Shifts: An Empirical Study with Visual Data
- Faster Improvement Rate Population Based Training
- A Simple Dynamic Learning Rate Tuning Algorithm For Automated Training of DNNs
- Deep learning for pedestrians: backpropagation in CNNs
- MG-WFBP: Efficient Data Communication for Distributed Synchronous SGD Algorithms
- Collage Inference: Using Coded Redundancy for Low Variance Distributed Image Classification
- XDA: Accurate, Robust Disassembly with Transfer Learning
- Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
- BS-NAS: Broadening-and-Shrinking One-Shot NAS with Searchable Numbers of Channels
- Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
- A New Look at Ghost Normalization
- Mask Guided Matting via Progressive Refinement Network
- Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent
- 1st place solution for AVA-Kinetics Crossover in AcitivityNet Challenge 2020
- Speeding up Deep Model Training by Sharing Weights and Then Unsharing
- Hydra: A Peer to Peer Distributed Training & Data Collection Framework
- Learning an Efficient Network for Large-Scale Hierarchical Object Detection with Data Imbalance: 3rd Place Solution to Open Images Challenge 2019
- DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers
- Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD
- Evaluating Transformers for Lightweight Action Recognition
- Fast Training of Sparse Graph Neural Networks on Dense Hardware
- What Happens after SGD Reaches Zero Loss? --A Mathematical Framework
- Stage-based Hyper-parameter Optimization for Deep Learning
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems
- High-Capacity Expert Binary Networks
- Distributed Low Precision Training Without Mixed Precision
- LIP: Local Importance-based Pooling
- Fast, Better Training Trick -- Random Gradient
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- On Large-Cohort Training for Federated Learning
- Salient Object Ranking with Position-Preserved Attention
- Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails
- Improving the convergence of SGD through adaptive batch sizes
- Why flatness does and does not correlate with generalization for deep neural networks
- Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs
- Dynamic Scheduling of MPI-based Distributed Deep Learning Training Jobs
- Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- Defending Against Image Corruptions Through Adversarial Augmentations
- CPM R-CNN: Calibrating Point-guided Misalignment in Object Detection
- Contrastive Attraction and Contrastive Repulsion for Representation Learning
- NetAdaptV2: Efficient Neural Architecture Search with Fast Super-Network Training and Architecture Optimization
- Fully-Automated Liver Tumor Localization and Characterization from Multi-Phase MR Volumes Using Key-Slice ROI Parsing: A Physician-Inspired Approach
- HALCONE : A Hardware-Level Timestamp-based Cache Coherence Scheme for Multi-GPU systems
- Joint-DetNAS: Upgrade Your Detector with NAS, Pruning and Dynamic Distillation
- Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
- Disentangling Sampling and Labeling Bias for Learning in Large-Output Spaces
- 3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
- Heterogeneity-aware Gradient Coding for Straggler Tolerance
- Exemplar-Based Open-Set Panoptic Segmentation Network
- SAFCAR: Structured Attention Fusion for Compositional Action Recognition
- Efficient Contextual Representation Learning Without Softmax Layer
- Accelerating Gossip SGD with Periodic Global Averaging
- Stochastic Reweighted Gradient Descent
- swCaffe: a Parallel Framework for Accelerating Deep Learning Applications on Sunway TaihuLight
- Large Batch Simulation for Deep Reinforcement Learning
- Content-Aware Preserving Image Generation
- Dynamic Collective Intelligence Learning: Finding Efficient Sparse Model via Refined Gradients for Pruned Weights
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
- Distributed Deep Learning Strategies For Automatic Speech Recognition
- NAS-HPO-Bench-II: A Benchmark Dataset on Joint Optimization of Convolutional Neural Network Architecture and Training Hyperparameters
- Distributed Reinforcement Learning of Targeted Grasping with Active Vision for Mobile Manipulators
- Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
- Improved generalization by noise enhancement
- Parameter Prediction for Unseen Deep Architectures
- Inefficiency of K-FAC for Large Batch Size Training
- Learning Temporally Invariant and Localizable Features via Data Augmentation for Video Recognition
- Nonlinear Conjugate Gradients For Scaling Synchronous Distributed DNN Training
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy Compression
- Scalable and Practical Natural Gradient for Large-Scale Deep Learning
- New Advances in Body Composition Assessment with ShapedNet: A Single Image Deep Regression Approach
- Do Self-Supervised and Supervised Methods Learn Similar Visual Representations?
- A Baseline Framework for Part-level Action Parsing and Action Recognition
- Drill the Cork of Information Bottleneck by Inputting the Most Important Data
- Long-Short Temporal Contrastive Learning of Video Transformers
- Concurrent Adversarial Learning for Large-Batch Training
- Probabilistic Ranking-Aware Ensembles for Enhanced Object Detections
- Highly Efficient Knowledge Graph Embedding Learning with Orthogonal Procrustes Analysis
- Asynchronous Distributed Optimization with Redundancy in Cost Functions
- Variance Reduction on General Adaptive Stochastic Mirror Descent
- Reverse engineering learned optimizers reveals known and novel mechanisms
- Representation Learning for Remote Sensing: An Unsupervised Sensor Fusion Approach
- Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning
- LIGAR: Lightweight General-purpose Action Recognition
- The Minimax Complexity of Distributed Optimization
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
- Adaptive Differentially Private Empirical Risk Minimization
- Controllable Orthogonalization in Training DNNs
- Distributed Deep Reinforcement Learning: An Overview
- Top-1 Solution of Multi-Moments in Time Challenge 2019
- Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
- Comparing the costs of abstraction for DL frameworks
- Balance-Oriented Focal Loss with Linear Scheduling for Anchor Free Object Detection
- Contrastive Semi-supervised Learning for ASR
- Quality-Aware Network for Human Parsing
- FixNorm: Dissecting Weight Decay for Training Deep Neural Networks
- Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
- Understanding the Effects of Pre-Training for Object Detectors via Eigenspectrum
- Minibatch Processing in Spiking Neural Networks
- Heterogeneous CPU+GPU Stochastic Gradient Descent Algorithms
- Linear Context Transform Block
- Instance Scale Normalization for image understanding
- Benchmark Tests of Convolutional Neural Network and Graph Convolutional Network on HorovodRunner Enabled Spark Clusters
- Map Generation from Large Scale Incomplete and Inaccurate Data Labels
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- Multi-step Estimation for Gradient-based Meta-learning
- Video Modeling with Correlation Networks
- Extrapolation for Large-batch Training in Deep Learning
- Hippo: Taming Hyper-parameter Optimization of Deep Learning with Stage Trees
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with Guarantees
- Accelerated Sparsified SGD with Error Feedback
- Question Guided Modular Routing Networks for Visual Question Answering
- Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
- Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
- A Survey on Large-scale Machine Learning
- Margin-Based Regularization and Selective Sampling in Deep Neural Networks
- Parsing R-CNN for Instance-Level Human Analysis
- Importance Weighted Evolution Strategies
- A Simple Non-i.i.d. Sampling Approach for Efficient Training and Better Generalization
- Selective sampling for accelerating training of deep neural networks
- Affine Self Convolution
- Large Batch Training Does Not Need Warmup
- Stochastic natural gradient descent draws posterior samples in function space
- Exploring Weight Symmetry in Deep Neural Networks
- Trajectory Normalized Gradients for Distributed Optimization
- Accelerating Training of Deep Neural Networks with a Standardization Loss
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
- SWIFT: Expedited Failure Recovery for Large-scale DNN Training
- Online Evolutionary Batch Size Orchestration for Scheduling Deep Learning Workloads in GPU Clusters
- Token Shift Transformer for Video Classification
- HyperSched: Dynamic Resource Reallocation for Model Development on a Deadline
- Scaling Distributed Training of Flood-Filling Networks on HPC Infrastructure for Brain Mapping
- On the Structural Sensitivity of Deep Convolutional Networks to the Directions of Fourier Basis Functions
- A Technical Report for VIPriors Image Classification Challenge
- How do SGD hyperparameters in natural training affect adversarial robustness?
- DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging
- Dynamic Graph: Learning Instance-aware Connectivity for Neural Networks
- SQWA: Stochastic Quantized Weight Averaging for Improving the Generalization Capability of Low-Precision Deep Neural Networks
- CSER: Communication-efficient SGD with Error Reset
- Improving compute efficacy frontiers with SliceOut
- Improving Efficiency in Large-Scale Decentralized Distributed Training
- False Negative Distillation and Contrastive Learning for Personalized Outfit Recommendation
- BFTrainer: Low-Cost Training of Neural Networks on Unfillable Supercomputer Nodes
- NG+ : A Multi-Step Matrix-Product Natural Gradient Method for Deep Learning
- Reducing the feature divergence of RGB and near-infrared images using Switchable Normalization
- E2-Train: Training State-of-the-art CNNs with Over 80% Energy Savings
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
- BPPSA: Scaling Back-propagation by Parallel Scan Algorithm
- Associating Multi-Scale Receptive Fields for Fine-grained Recognition
- Self-Distilled Self-Supervised Representation Learning
- SURREAL-System: Fully-Integrated Stack for Distributed Deep Reinforcement Learning
- Deep set conditioned latent representations for action recognition
- FOSNet: An End-to-End Trainable Deep Neural Network for Scene Recognition
- Scaling Distributed Training with Adaptive Summation
- Gradient Sparification for Asynchronous Distributed Training
- SSAN: Separable Self-Attention Network for Video Representation Learning
- Finite-Time Consensus Learning for Decentralized Optimization with Nonlinear Gossiping
- Data-parallel distributed training of very large models beyond GPU capacity
- Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
- AccUDNN: A GPU Memory Efficient Accelerator for Training Ultra-deep Neural Networks
- Self-Reorganizing and Rejuvenating CNNs for Increasing Model Capacity Utilization
- Contrastive Weight Regularization for Large Minibatch SGD
- Distributed Sparse SGD with Majority Voting
- Resource-Adaptive Successive Doubling for Hyperparameter Optimization with Large Datasets on High-Performance Computing Systems
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Beyond the Memory Wall: A Case for Memory-centric HPC System for Deep Learning
- Joint Matrix Decomposition for Deep Convolutional Neural Networks Compression
- High Throughput Synchronous Distributed Stochastic Gradient Descent
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- Toward Efficient Online Scheduling for Distributed Machine Learning Systems
- Batch Normalization Sampling
- Error Compensated Loopless SVRG, Quartz, and SDCA for Distributed Optimization
- Enhance Diffusion to Improve Robust Generalization
- Parallelizable Stack Long Short-Term Memory
- Revisiting Explicit Regularization in Neural Networks for Well-Calibrated Predictive Uncertainty
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Understanding the Disharmony between Weight Normalization Family and Weight Decay: shifted Regularizer
- Scalable Smartphone Cluster for Deep Learning
- Asynchronous Decentralized Distributed Training of Acoustic Models
- HydaLearn: Highly Dynamic Task Weighting for Multi-task Learning with Auxiliary Tasks
- OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- Gap Aware Mitigation of Gradient Staleness
- Writing in The Air: Unconstrained Text Recognition from Finger Movement Using Spatio-Temporal Convolution
- A Distributed SGD Algorithm with Global Sketching for Deep Learning Training Acceleration
- Elastic deep learning in multi-tenant GPU cluster
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- Critical Parameters for Scalable Distributed Learning with Large Batches and Asynchronous Updates
- Learning degraded image classification with restoration data fidelity
- Learning a Unified Embedding for Visual Search at Pinterest
- A Generalization of the Allreduce Operation
- How Data Augmentation affects Optimization for Linear Regression
- Why Do Better Loss Functions Lead to Less Transferable Features?
- DEED: A General Quantization Scheme for Communication Efficiency in Bits
- Mini-batch Serialization: CNN Training with Inter-layer Data Reuse
- Genetic-algorithm-optimized neural networks for gravitational wave classification
- A Predictive Autoscaler for Elastic Batch Jobs
- Dynamic Slimmable Network
- Progressive Multi-stage Feature Mix for Person Re-Identification
- PGT: A Progressive Method for Training Models on Long Videos
- Improving Differentially Private SGD via Randomly Sparsified Gradients
- Secure Distributed Training at Scale
- High Throughput Training of Deep Surrogates from Large Ensemble Runs
- Quality-Aware Network for Face Parsing
- Towards Explainable Fact Checking
- Weighted Aggregating Stochastic Gradient Descent for Parallel Deep Learning
- Addressing Algorithmic Bottlenecks in Elastic Machine Learning with Chicle
- POD: Practical Object Detection with Scale-Sensitive Network
- LocalNorm: Robust Image Classification through Dynamically Regularized Normalization
- Ouroboros: On Accelerating Training of Transformer-Based Language Models
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- Spatially Attentive Output Layer for Image Classification
- IB-DRR: Incremental Learning with Information-Back Discrete Representation Replay
- GoG: Relation-aware Graph-over-Graph Network for Visual Dialog
- Implicit Regularization of Bregman Proximal Point Algorithm and Mirror Descent on Separable Data
- High performance and energy efficient inference for deep learning on ARM processors
- Slim-DP: A Light Communication Data Parallelism for DNN
- Learning Metrics from Mean Teacher: A Supervised Learning Method for Improving the Generalization of Speaker Verification System
- Towards Better Generalization: BP-SVRG in Training Deep Neural Networks
- How to Train your DNN: The Network Operator Edition
- Caramel: Accelerating Decentralized Distributed Deep Learning with Computation Scheduling
- PointManifoldCut: Point-wise Augmentation in the Manifold for Point Clouds
- Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories
- Progressive Compressed Records: Taking a Byte out of Deep Learning Data
- Push for Center Learning via Orthogonalization and Subspace Masking for Person Re-Identification
- AutoCLINT: The Winning Method in AutoCV Challenge 2019
- Scaling the training of particle classification on simulated MicroBooNE events to multiple GPUs
- MG-WFBP: Merging Gradients Wisely for Efficient Communication in Distributed Deep Learning
- Metastatic Cancer Image Classification Based On Deep Learning Method
- ISyNet: Convolutional Neural Networks design for AI accelerator
- Oscars: Adaptive Semi-Synchronous Parallel Model for Distributed Deep Learning with Global View
- Momentum Improves Normalized SGD
- Is High Variance Unavoidable in RL? A Case Study in Continuous Control
- Exploiting Redundancy: Separable Group Convolutional Networks on Lie Groups
- Analyzing the benefits of communication channels between deep learning models
- Batch size-invariance for policy optimization
- Towards Enhancing Fine-grained Details for Image Matting
- AsymptoticNG: A regularized natural gradient optimization algorithm with look-ahead strategy
- Hoplite: Efficient and Fault-Tolerant Collective Communication for Task-Based Distributed Systems
- Cyclic orthogonal convolutions for long-range integration of features
- GenURL: A General Framework for Unsupervised Representation Learning
- Discretization-Aware Architecture Search
- Accumulated Decoupled Learning: Mitigating Gradient Staleness in Inter-Layer Model Parallelization
- Kernelized Classification in Deep Networks
- Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
- Improving Layer-wise Adaptive Rate Methods using Trust Ratio Clipping
- Data-driven Algorithm Selection and Parameter Tuning: Two Case studies in Optimization and Signal Processing
- On Compositions of Transformations in Contrastive Self-Supervised Learning
- Effective Approaches to Batch Parallelization for Dynamic Neural Network Architectures
- Stochastic Proximal Gradient Algorithm with Minibatches. Application to Large Scale Learning Models
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- PatchGame: Learning to Signal Mid-level Patches in Referential Games
- To Talk or to Work: Delay Efficient Federated Learning over Mobile Edge Devices
- Relation/Entity-Centric Reading Comprehension
- Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods
- Trade-offs of Local SGD at Scale: An Empirical Study
- Integrated Model, Batch and Domain Parallelism in Training Neural Networks
- Procrustes: a Dataflow and Accelerator for Sparse Deep Neural Network Training
- The Impact of Spatiotemporal Augmentations on Self-Supervised Audiovisual Representation Learning
- MURAUER: Mapping Unlabeled Real Data for Label AUstERity
- Knowledge Fusion Transformers for Video Action Recognition
- Farkas layers: don't shift the data, fix the geometry
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-Adaptation
- Task allocation for decentralized training in heterogeneous environment
- An Empirical Study on Compressed Decentralized Stochastic Gradient Algorithms with Overparameterized Models
- Gradient-based Hyperparameter Optimization Over Long Horizons
- Influence-Balanced Loss for Imbalanced Visual Classification
- Efficient Modelling Across Time of Human Actions and Interactions
- A Technical Report for ICCV 2021 VIPriors Re-identification Challenge
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
- Evaluating the fairness of fine-tuning strategies in self-supervised learning
- Frequency-aware SGD for Efficient Embedding Learning with Provable Benefits
- A Spark ML driven preprocessing approach for deep learning based scholarly data applications
- Exploring and Improving Mobile Level Vision Transformers
- : Accelerating Asynchronous Communication in Decentralized Deep Learning
- 4-bit Quantization of LSTM-based Speech Recognition Models
- Stochastic Contrastive Learning
- Decision Propagation Networks for Image Classification
- Logit Attenuating Weight Normalization
- AutoBSS: An Efficient Algorithm for Block Stacking Style Search
- Tensor Yard: One-Shot Algorithm of Hardware-Friendly Tensor-Train Decomposition for Convolutional Neural Networks
- Pipelined Training with Stale Weights of Deep Convolutional Neural Networks
- Domain Adaptor Networks for Hyperspectral Image Recognition
- Label Denoising with Large Ensembles of Heterogeneous Neural Networks
- Joint Learning of Instance and Semantic Segmentation for Robotic Pick-and-Place with Heavy Occlusions in Clutter
- AAG: Self-Supervised Representation Learning by Auxiliary Augmentation with GNT-Xent Loss
- JUWELS Booster -- A Supercomputer for Large-Scale AI Research
- Accelerating Distributed K-FAC with Smart Parallelism of Computing and Communication Tasks
- A Highly Efficient Distributed Deep Learning System For Automatic Speech Recognition
- Poly-NL: Linear Complexity Non-local Layers with Polynomials
- Incorporating NODE with Pre-trained Neural Differential Operator for Learning Dynamics
- Kanerva++: extending The Kanerva Machine with differentiable, locally block allocated latent memory
- USB: Universal-Scale Object Detection Benchmark
- The Gradient Convergence Bound of Federated Multi-Agent Reinforcement Learning with Efficient Communication
- LRTuner: A Learning Rate Tuner for Deep Neural Networks
- Cloud Collectives: Towards Cloud-aware Collectives forML Workloads with Rank Reordering
- A Sum-of-Ratios Multi-Dimensional-Knapsack Decomposition for DNN Resource Scheduling
- Estimating the Uncertainty of Neural Network Forecasts for Influenza Prevalence Using Web Search Activity
- Approximate Random Dropout
- CITIES: Contextual Inference of Tail-Item Embeddings for Sequential Recommendation