Augmenting Supervised Neural Networks with Unsupervised Objectives for Large-scale Image Classification
arXiv:1606.06582
Abstract
Unsupervised learning and supervised learning are key research topics in deep learning. However, as high-capacity supervised neural networks trained with a large amount of labels have achieved remarkable success in many computer vision tasks, the availability of large-scale labeled images reduced the significance of unsupervised learning. Inspired by the recent trend toward revisiting the importance of unsupervised learning, we investigate joint supervised and unsupervised learning in a large-scale setting by augmenting existing neural networks with decoding pathways for reconstruction. First, we demonstrate that the intermediate activations of pretrained large-scale classification networks preserve almost all the information of input images except a portion of local spatial details. Then, by end-to-end training of the entire augmented architecture with the reconstructive objective, we show improvement of the network performance for supervised tasks. We evaluate several variants of autoencoders, including the recently proposed "what-where" autoencoder that uses the encoder pooling switches, to study the importance of the architecture design. Taking the 16-layer VGGNet trained under the ImageNet ILSVRC 2012 protocol as a strong baseline for image classification, our methods improve the validation-set accuracy by a noticeable margin.
International Conference on Machine Learning (ICML), 2016
References in corpus (4)
Cited by in corpus (16)
- Adapting Auxiliary Losses Using Gradient Similarity
- ExtremeWeather: A large-scale climate dataset for semi-supervised detection, localization, and understanding of extreme weather events
- Robust Conditional Generative Adversarial Networks
- Pre-training with Non-expert Human Demonstration for Deep Reinforcement Learning
- Hebbian Semi-Supervised Learning in a Sample Efficiency Setting
- Classification-Reconstruction Learning for Open-Set Recognition
- Semi-Supervised Phoneme Recognition with Recurrent Ladder Networks
- Scaling Object Detection by Transferring Classification Weights
- Deeply Supervised Semantic Model for Click-Through Rate Prediction in Sponsored Search
- Contextual Classification Using Self-Supervised Auxiliary Models for Deep Neural Networks
- Deep Global-Connected Net With The Generalized Multi-Piecewise ReLU Activation in Deep Learning
- Text-to-image Synthesis via Symmetrical Distillation Networks
- Training Multimodal Systems for Classification with Multiple Objectives
- Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers
- Learning Rich Representations For Structured Visual Prediction Tasks
- Image Embedded Segmentation: Uniting Supervised and Unsupervised Objectives for Segmenting Histopathological Images