Efficient Classification of Very Large Images with Tiny Objects
arXiv:2106.02694
Abstract
An increasing number of applications in computer vision, specially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: ) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources.
References in corpus (14)
- Understanding deep learning requires rethinking generalization
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Co-teaching: Robust Training of Deep Neural Networks with Extremely Noisy Labels
- A Closer Look at Memorization in Deep Networks
- Neural Image Compression for Gigapixel Histopathology Image Analysis
- Three Factors Influencing Minima in SGD
- Supervised Contrastive Learning
- Streaming convolutional neural networks for end-to-end learning with multi-megapixel images
- Classification and Disease Localization in Histopathology Using Only Global Labels: A Weakly-Supervised Approach
- Global Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification
- Processing Megapixel Images with Deep Attention-Sampling Models
- Hard-Attention for Scalable Image Classification
- Needles in Haystacks: On Classifying Tiny Objects in Large Images
- Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects