A robust estimator of mutual information for deep learning interpretability
arXiv:2211.00024 · doi:10.1088/2632-2153/acc444
Abstract
We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced Jimmie), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established mutual information estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train deep learning models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available.
30 pages, 8 figures. Minor changes to match version accepted for publication in Machine Learning: Science and Technology. GMM-MI available at https://github.com/dpiras/GMM-MI
References in corpus (8)
- Likelihood-free inference with neural compression of DES SV weak lensing map statistics
- Mutual Information as a Tool for Identifying Phase Transitions in Dynamical Complex Systems With Limited Data
- How much a galaxy knows about its large-scale environment?: An information theoretic perspective
- Symmetries and phase diagrams with real-space mutual information neural estimation
- Machines Learn to Infer Stellar Parameters Just by Looking at a Large Number of Spectra
- Sufficiency of a Gaussian power spectrum likelihood for accurate cosmology from upcoming weak lensing surveys
- Do galactic bars depend on environment?: An information theoretic analysis of Galaxy Zoo 2
- On the origin of red spirals: Does assembly bias play a role?
Cited by in corpus (9)
- Data Compression and Inference in Cosmology with Self-Supervised Machine Learning
- Explaining dark matter halo density profiles with neural networks
- Deep learning insights into cosmological structure formation
- Cosmological feedback from a halo assembly perspective
- A representation learning approach to probe for dynamical dark energy in matter power spectra
- Deep learning insights into non-universality in the halo mass function
- Data-Space Validation of High-Dimensional Models by Comparing Sample Quantiles
- CDM and early dark energy in latent space: a data-driven parametrization of the CMB temperature power spectrum
- QUEST (Quasar Unsupervised Encoder and Synthesis Tool): A machine learning framework to generate quasar spectra