Learning Discrete Representations via Information Maximizing Self-Augmented Training
arXiv:1702.08720
Abstract
Learning discrete representations of data is a central machine learning task because of the compactness of the representations and ease of interpretation. The task includes clustering and hash learning as special cases. Deep neural networks are promising to be used because they can model the non-linearity of data and scale to large datasets. However, their model complexity is huge, and therefore, we need to carefully regularize the networks in order to learn useful representations that exhibit intended invariance for applications of interest. To this end, we propose a method called Information Maximizing Self-Augmented Training (IMSAT). In IMSAT, we use data augmentation to impose the invariance on discrete representations. More specifically, we encourage the predicted representations of augmented data points to be close to those of the original data points in an end-to-end fashion. At the same time, we maximize the information-theoretic dependency between data and their predicted discrete representations. Extensive experiments on benchmark datasets show that IMSAT produces state-of-the-art results for both clustering and unsupervised hash learning.
To appear at ICML 2017
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improving neural networks by preventing co-adaptation of feature detectors
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders
- Learning with Pseudo-Ensembles
Cited by in corpus (88)
- Unsupervised Data Augmentation for Consistency Training
- On Mutual Information Maximization for Representation Learning
- A Comprehensive Survey on Test-Time Adaptation under Distribution Shifts
- A survey on Semi-, Self- and Unsupervised Learning for Image Classification
- Clustering with Deep Learning: Taxonomy and New Methods
- Twin Contrastive Learning for Online Clustering
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation
- SpectralNet: Spectral Clustering using Deep Neural Networks
- Joint Optimization Framework for Learning with Noisy Labels
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- Labelling unlabelled videos from scratch with multi-modal self-supervision
- Putting An End to End-to-End: Gradient-Isolated Learning of Representations
- ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction
- Transductive Information Maximization For Few-Shot Learning
- On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data
- Combining pretrained CNN feature extractors to enhance clustering of complex natural images
- Gaussian Mixture Generative Adversarial Networks for Diverse Datasets, and the Unsupervised Clustering of Images
- Clustering-friendly Representation Learning via Instance Discrimination and Feature Decorrelation
- Stacked Capsule Autoencoders
- AVA-AVD: Audio-Visual Speaker Diarization in the Wild
- Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach
- Information Maximization Clustering via Multi-View Self-Labelling
- TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation
- Source Data-absent Unsupervised Domain Adaptation through Hypothesis Transfer and Labeling Transfer
- Transformer-Based Source-Free Domain Adaptation
- Image Clustering using an Augmented Generative Adversarial Network and Information Maximization
- Unsupervised Semantic Aggregation and Deformable Template Matching for Semi-Supervised Learning
- Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization
- RDEC: Integrating Regularization into Deep Embedded Clustering for Imbalanced Datasets
- AnchorGAE: General Data Clustering via Bipartite Graph Convolution
- An Explicit Local and Global Representation Disentanglement Framework with Applications in Deep Clustering and Unsupervised Object Detection
- Improving Image Clustering With Multiple Pretrained CNN Feature Extractors
- Elastic-InfoGAN: Unsupervised Disentangled Representation Learning in Class-Imbalanced Data
- Classification from Pairwise Similarity and Unlabeled Data
- ConDA: Continual Unsupervised Domain Adaptation
- Doubly Contrastive Deep Clustering
- MiCE: Mixture of Contrastive Experts for Unsupervised Image Clustering
- GATCluster: Self-Supervised Gaussian-Attention Network for Image Clustering
- M2IOSR: Maximal Mutual Information Open Set Recognition
- SubTab: Subsetting Features of Tabular Data for Self-Supervised Representation Learning
- Improving Unsupervised Image Clustering With Robust Learning
- PointSmile: Point Self-supervised Learning via Curriculum Mutual Information
- Learning Domain Invariant Representations by Joint Wasserstein Distance Minimization
- Evolution Is All You Need: Phylogenetic Augmentation for Contrastive Learning
- Improving ClusterGAN Using Self-Augmented Information Maximization of Disentangling Latent Spaces
- DINE: Domain Adaptation from Single and Multiple Black-box Predictors
- Self-Supervised Learning by Estimating Twin Class Distributions
- Self-labeled Conditional GANs
- Faster Convergence in Deep-Predictive-Coding Networks to Learn Deeper Representations
- Multi-Facet Clustering Variational Autoencoders
- On Evolving Attention Towards Domain Adaptation
- Consistency Regularization for Cross-Lingual Fine-Tuning
- Meta-Learning to Cluster
- On the Estimation of Information Measures of Continuous Distributions
- Deep Fair Discriminative Clustering
- DRo: A data-scarce mechanism to revolutionize the performance of Deep Learning based Security Systems
- Focus of Attention Improves Information Transfer in Visual Features
- Disentangling to Cluster: Gaussian Mixture Variational Ladder Autoencoders
- Source-Free Domain Adaptation for Question Answering with Masked Self-training
- Infomax Neural Joint Source-Channel Coding via Adversarial Bit Flip
- Topological Gradient-based Competitive Learning
- Dissimilarity Mixture Autoencoder for Deep Clustering
- Deep Inverse Feature Learning: A Representation Learning of Error
- Augmented Cyclic Consistency Regularization for Unpaired Image-to-Image Translation
- Self-Supervised Domain Adaptation with Consistency Training
- Simple, Scalable, and Stable Variational Deep Clustering
- Hypothesis Disparity Regularized Mutual Information Maximization
- MUSCLE: Strengthening Semi-Supervised Learning Via Concurrent Unsupervised Learning Using Mutual Information Maximization
- Integrating Categorical Semantics into Unsupervised Domain Translation
- Limited Gradient Descent: Learning With Noisy Labels
- Double cycle-consistent generative adversarial network for unsupervised conditional generation
- Neural Bayes: A Generic Parameterization Method for Unsupervised Representation Learning
- You Never Cluster Alone
- Diversified Multi-prototype Representation for Semi-supervised Segmentation
- Improving Image Clustering through Sample Ranking and Its Application to remote--sensing images
- Deep Transformation-Invariant Clustering
- A Framework for Deep Constrained Clustering
- Multi-level Feature Learning on Embedding Layer of Convolutional Autoencoders and Deep Inverse Feature Learning for Image Clustering
- Learning Interpretable and Discrete Representations with Adversarial Training for Unsupervised Text Classification
- Deep Kernel Learning for Clustering
- Unsupervised Image Segmentation by Mutual Information Maximization and Adversarial Regularization
- A Comparison of Discrete Latent Variable Models for Speech Representation Learning
- Contrastive Representation Learning with Trainable Augmentation Channel
- Adversarial confidence and smoothness regularizations for scalable unsupervised discriminative learning
- Generalized Clustering by Learning to Optimize Expected Normalized Cuts
- Unsupervised Representation Learning via Neural Activation Coding
- Think Global, Act Local: Relating DNN generalisation and node-level SNR
- Deep Clustering with Measure Propagation