MINE: Mutual Information Neural Estimation
arXiv:1801.04062
Abstract
We argue that the estimation of mutual information between high dimensional continuous random variables can be achieved by gradient descent over neural networks. We present a Mutual Information Neural Estimator (MINE) that is linearly scalable in dimensionality as well as in sample size, trainable through back-prop, and strongly consistent. We present a handful of applications on which MINE can be used to minimize or maximize mutual information. We apply MINE to improve adversarially trained generative models. We also use MINE to implement Information Bottleneck, applying it to supervised classification; our results demonstrate substantial improvement in flexibility and performance in these settings.
19 pages, 6 figures
Cited by in corpus (89)
- What Makes for Good Views for Contrastive Learning?
- Contrastive Multiview Coding
- Learning Video Representations using Contrastive Bidirectional Transformer
- HDMI: High-order Deep Multiplex Infomax
- Entropy and mutual information in models of deep neural networks
- Isolating Sources of Disentanglement in Variational Autoencoders
- FMix: Enhancing Mixed Sample Data Augmentation
- Knowledge-Preserving Incremental Social Event Detection via Heterogeneous GNNs
- Learning Robust Representations via Multi-View Information Bottleneck
- A Mutual Information Maximization Perspective of Language Representation Learning
- Revisiting Training Strategies and Generalization Performance in Deep Metric Learning
- DisCo Fever: Robust Networks Through Distance Correlation
- Multimodal Emotion Recognition Using Deep Canonical Correlation Analysis
- Heterogeneous Deep Graph Infomax
- Formal Limitations on the Measurement of Mutual Information
- Near-Optimal Representation Learning for Hierarchical Reinforcement Learning
- LOGAN: Latent Optimisation for Generative Adversarial Networks
- Graph Clustering with Graph Neural Networks
- Graph Pooling via Coarsened Graph Infomax
- Graph Representation Learning via Graphical Mutual Information Maximization
- Normalization of breast MRIs using Cycle-Consistent Generative Adversarial Networks
- F-BLEAU: Fast Black-box Leakage Estimation
- Understanding the Limitations of Variational Mutual Information Estimators
- A Tutorial on Deep Latent Variable Models of Natural Language
- Extracting Parallel Sentences with Bidirectional Recurrent Neural Networks to Improve Machine Translation
- Feature Overcorrelation in Deep Graph Neural Networks: A New Perspective
- Deep Reinforcement and InfoMax Learning
- Semantic Information G Theory and Logical Bayesian Inference for Machine Learning
- Boosting Semi-supervised Image Segmentation with Global and Local Mutual Information Regularization
- Self-Supervised Learning with Kernel Dependence Maximization
- Mutual Information Scaling for Tensor Network Machine Learning
- Estimating the Mutual Information between two Discrete, Asymmetric Variables with Limited Samples
- Unsupervised Attributed Multiplex Network Embedding
- Conditional Mutual Information Neural Estimator
- Caveats for information bottleneck in deterministic scenarios
- Adaptive Deep Kernel Learning
- Incorporating Attributes and Multi-Scale Structures for Heterogeneous Graph Contrastive Learning
- Boosting Discriminative Visual Representation Learning with Scenario-Agnostic Mixup
- Variational Bayesian Optimal Experimental Design with Normalizing Flows
- Learning Robust and Multilingual Speech Representations
- Modeling Multiple Views via Implicitly Preserving Global Consistency and Local Complementarity
- Demystifying Deep Learning in Predictive Spatio-Temporal Analytics: An Information-Theoretic Framework
- Self-supervised Learning from a Multi-view Perspective
- Information-Theoretic Text Hallucination Reduction for Video-grounded Dialogue
- M2IOSR: Maximal Mutual Information Open Set Recognition
- MemDA: Forecasting Urban Time Series with Memory-based Drift Adaptation
- Neural Estimators for Conditional Mutual Information Using Nearest Neighbors Sampling
- Multi-label Contrastive Predictive Coding
- Information Losses in Neural Classifiers from Sampling
- Fairness-Aware Neural Réyni Minimization for Continuous Features
- Learning Adversarially Robust Representations via Worst-Case Mutual Information Maximization
- To Split or Not to Split: The Impact of Disparate Treatment in Classification
- Mutual Information Gradient Estimation for Representation Learning
- VFunc: a Deep Generative Model for Functions
- Learning Actionable Representations with Goal-Conditioned Policies
- QAInfomax: Learning Robust Question Answering System by Mutual Information Maximization
- Dynamic learning rate using Mutual Information
- Reviewing Evolution of Learning Functions and Semantic Information Measures for Understanding Deep Learning
- Training Normalizing Flows with the Information Bottleneck for Competitive Generative Classification
- InfoShape: Task-Based Neural Data Shaping via Mutual Information
- Neural Methods for Point-wise Dependency Estimation
- Biphasic Learning of GANs for High-Resolution Image-to-Image Translation
- DHOG: Deep Hierarchical Object Grouping
- Improving the Learning of Multi-column Convolutional Neural Network for Crowd Counting
- Empowerment-driven Exploration using Mutual Information Estimation
- Mutual Information State Intrinsic Control
- Regularity Normalization: Neuroscience-Inspired Unsupervised Attention across Neural Network Layers
- General Probabilistic Surface Optimization and Log Density Estimation
- MIM: Mutual Information Machine
- Learning Segmentation Masks with the Independence Prior
- Mitigating the Effects of Non-Identifiability on Inference for Bayesian Neural Networks with Latent Variables
- Deep Music Information Dynamics
- SentenceMIM: A Latent Variable Language Model
- A Quantitative Metric for Privacy Leakage in Federated Learning
- Non-Local Representation based Mutual Affine-Transfer Network for Photorealistic Stylization
- Mutual Information Maximization in Graph Neural Networks
- Better Long-Range Dependency By Bootstrapping A Mutual Information Regularizer
- Graph Embedding Using Infomax for ASD Classification and Brain Functional Difference Detection
- Learning Deep Representations by Mutual Information for Person Re-identification
- Exploiting Style and Attention in Real-World Super-Resolution
- Toward Interpretability of Dual-Encoder Models for Dialogue Response Suggestions
- Information Theory-Guided Heuristic Progressive Multi-View Coding
- Modeling Psychotherapy Dialogues with Kernelized Hashcode Representations: A Nonparametric Information-Theoretic Approach
- On the Veracity of Cyber Intrusion Alerts Synthesized by Generative Adversarial Networks
- TzK: Flow-Based Conditional Generative Model
- High Mutual Information in Representation Learning with Symmetric Variational Inference
- Adversarial Orthogonal Regression: Two non-Linear Regressions for Causal Inference
- Informative GANs via Structured Regularization of Optimal Transport
- A Novel Information-Theoretic Objective to Disentangle Representations for Fair Classification