Improving Deep Image Clustering With Spatial Transformer Layers
arXiv:1902.05401 · doi:10.1007/978-3-030-30490-4_51
Abstract
Image clustering is an important but challenging task in machine learning. As in most image processing areas, the latest improvements came from models based on the deep learning approach. However, classical deep learning methods have problems to deal with spatial image transformations like scale and rotation. In this paper, we propose the use of visual attention techniques to reduce this problem in image clustering methods. We evaluate the combination of a deep image clustering model called Deep Adaptive Clustering (DAC) with the Spatial Transformer Networks (STN). The proposed model is evaluated in the datasets MNIST and FashionMNIST and outperformed the baseline model.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Striving for Simplicity: The All Convolutional Net
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- Towards K-means-friendly Spaces: Simultaneous Deep Learning and Clustering
- Variational Deep Embedding: An Unsupervised and Generative Approach to Clustering