Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
arXiv:1707.02968
Abstract
The success of deep learning in vision can be attributed to: (a) models with high capacity; (b) increased computational power; and (c) availability of large-scale labeled data. Since 2012, there have been significant advances in representation capabilities of the models and computational capabilities of GPUs. But the size of the biggest dataset has surprisingly remained constant. What will happen if we increase the dataset size by 10x or 100x? This paper takes a step towards clearing the clouds of mystery surrounding the relationship between `enormous data' and visual deep learning. By exploiting the JFT-300M dataset which has more than 375M noisy labels for 300M images, we investigate how the performance of current vision tasks would change if this data was used for representation learning. Our paper delivers some surprising (and some expected) findings. First, we find that the performance on vision tasks increases logarithmically based on volume of training data size. Second, we show that representation learning (or pre-training) still holds a lot of promise. One can improve performance on many vision tasks by just training a better base model. Finally, as expected, we present new state-of-the-art results for different vision tasks including image classification, object detection, semantic segmentation and human pose estimation. Our sincere hope is that this inspires vision community to not undervalue the data and develop collective efforts in building larger datasets.
ICCV 2017 camera ready
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Generating Videos with Scene Dynamics
- What makes ImageNet good for transfer learning?
- Analyzing the Performance of Multilayer Neural Networks for Object Recognition
- Large-Scale Deep Learning on the YFCC100M Dataset
Cited by in corpus (25)
- Snorkel: Rapid Training Data Creation with Weak Supervision
- Rethinking the Hyperparameters for Fine-tuning
- Data Distillation: Towards Omni-Supervised Learning
- Scale out for large minibatch SGD: Residual network training on ImageNet-1K with improved accuracy and reduced time to train
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Quantifying the Performance of Federated Transfer Learning
- Addressing Missing Labels in Large-Scale Sound Event Recognition Using a Teacher-Student Framework With Loss Masking
- Knowledge Concentration: Learning 100K Object Classifiers in a Single CNN
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- KINN: Incorporating Expert Knowledge in Neural Networks
- Share your Model instead of your Data: Privacy Preserving Mimic Learning for Ranking
- Tree-structured Kronecker Convolutional Network for Semantic Segmentation
- Stochastic Downsampling for Cost-Adjustable Inference and Improved Regularization in Convolutional Networks
- Robust and On-the-fly Dataset Denoising for Image Classification
- An empirical study of pretrained representations for few-shot classification
- To Pretrain or Not to Pretrain: Examining the Benefits of Pretraining on Resource Rich Tasks
- Designing for the Long Tail of Machine Learning
- Impact of Training Dataset Size on Neural Answer Selection Models
- Improving image generative models with human interactions
- Can We Achieve More with Less? Exploring Data Augmentation for Toxic Comment Classification
- Unsupervised Multi-label Dataset Generation from Web Data
- Dynamic Graph Correlation Learning for Disease Diagnosis with Incomplete Labels
- Classification of sparsely labeled spatio-temporal data through semi-supervised adversarial learning
- Function space analysis of deep learning representation layers
- Parameter Reference Loss for Unsupervised Domain Adaptation