Attending Category Disentangled Global Context for Image Classification
arXiv:1812.06663
Abstract
In this paper, we propose a general framework for image classification using the attention mechanism and global context, which could incorporate with various network architectures to improve their performance. To investigate the capability of the global context, we compare four mathematical models and observe the global context encoded in the category disentangled conditional generative model could give more guidance as "know what is task irrelevant will also know what is relevant". Based on this observation, we define a novel Category Disentangled Global Context (CDGC) and devise a deep network to obtain it. By attending CDGC, the baseline networks could identify the objects of interest more accurately, thus improving the performance. We apply the framework to many different network architectures and compare with the state-of-the-art on four publicly available datasets. Extensive results validate the effectiveness and superiority of our approach. Code will be made public upon paper acceptance.
I have not gotten legal permission to post it on arXiv from the corresponding authors
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improved Training of Wasserstein GANs
- Object Detectors Emerge in Deep Scene CNNs
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- Disentangling factors of variation in deep representations using adversarial training
- Discovering Hidden Factors of Variation in Deep Networks
- Learn To Pay Attention
- Label Tree Embeddings for Acoustic Scene Classification
- Watch, Listen, and Describe: Globally and Locally Aligned Cross-Modal Attentions for Video Captioning