Universal representations:The missing link between faces, text, planktons, and cat breeds
arXiv:1701.07275
Abstract
With the advent of large labelled datasets and high-capacity models, the performance of machine vision systems has been improving rapidly. However, the technology has still major limitations, starting from the fact that different vision problems are still solved by different models, trained from scratch or fine-tuned on the target data. The human visual system, in stark contrast, learns a universal representation for vision in the early life of an individual. This representation works well for an enormous variety of vision problems, with little or no change, with the major advantage of requiring little training data to solve any of them. In this paper we investigate whether neural networks may work as universal representations by studying their capacity in relation to the “size†of a large combination of vision problems. We do so by showing that a single neural network can learn simultaneously several very different visual domains (from sketches to planktons and MNIST digits) as well as, or better than, a number of specialized networks. However, we also show that this requires to carefully normalize the information in the network, by using domain-specific scaling factors or, more generically, by using an instance normalization layer.
10 pages, 4 figures, 5 tables
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Learning Transferable Features with Deep Adaptation Networks
- FitNets: Hints for Thin Deep Nets
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Return of the Devil in the Details: Delving Deep into Convolutional Nets
- Layer Normalization
- Deep convolutional filter banks for texture recognition and segmentation
Cited by in corpus (45)
- A Survey of Deep Learning-based Object Detection
- Multi-Task Learning for Dense Prediction Tasks: A Survey
- GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
- Learning to Generalize: Meta-Learning for Domain Generalization
- Meta-Learning with Warped Gradient Descent
- A Universal Representation Transformer Layer for Few-Shot Image Classification
- Contrastive Learning with Adversarial Examples
- Normalization Techniques in Training DNNs: Methodology, Analysis and Application
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- Depthwise Convolution is All You Need for Learning Multiple Visual Domains
- Learning to Optimize Domain Specific Normalization for Domain Generalization
- Episodic Training for Domain Generalization
- Generalizable multi-task, multi-domain deep segmentation of sparse pediatric imaging datasets via multi-scale contrastive regularization and multi-joint anatomical priors
- Beyond Shared Hierarchies: Deep Multitask Learning through Soft Layer Ordering
- Towards Universal Object Detection by Domain Attention
- Towards Universal Representation Learning for Deep Face Recognition
- Selecting Relevant Features from a Multi-domain Representation for Few-shot Classification
- Unpaired Multi-modal Segmentation via Knowledge Distillation
- Cross-domain Few-shot Learning with Task-specific Adapters
- Learning a Universal Template for Few-shot Dataset Generalization
- Generalization in multitask deep neural classifiers: a statistical physics approach
- 3D U-Net: A 3D Universal U-Net for Multi-Domain Medical Image Segmentation
- Style Normalization and Restitution for Domain Generalization and Adaptation
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Regularizing activations in neural networks via distribution matching with the Wasserstein metric
- A Lifelong Learning Approach to Brain MR Segmentation Across Scanners and Protocols
- Latent Domain Learning with Dynamic Residual Adapters
- Object Detection with a Unified Label Space from Multiple Datasets
- Disjoint Multi-task Learning between Heterogeneous Human-centric Tasks
- Learning Finer-class Networks for Universal Representations
- Label-efficient audio classification through multitask learning and self-supervision
- Deep Anomaly Detection by Residual Adaptation
- Text is Text, No Matter What: Unifying Text Recognition using Knowledge Distillation
- The Traveling Observer Model: Multi-task Learning Through Spatial Variable Embeddings
- Visually Grounded Continual Language Learning with Selective Specialization
- Multi-Task, Multi-Domain Deep Segmentation with Shared Representations and Contrastive Regularization for Sparse Pediatric Datasets
- Efficient Multi-Domain Network Learning by Covariance Normalization
- Introducing Pose Consistency and Warp-Alignment for Self-Supervised 6D Object Pose Estimation in Color Images
- CRL: Class Representative Learning for Image Classification
- What and Where: Learn to Plug Adapters via NAS for Multi-Domain Learning
- Multi-path Neural Networks for On-device Multi-domain Visual Classification
- Supervised Momentum Contrastive Learning for Few-Shot Classification
- Beyond without Forgetting: Multi-Task Learning for Classification with Disjoint Datasets
- Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain
- Disentangling Transfer and Interference in Multi-Domain Learning