Info-Clustering: A Mathematical Theory for Data Clustering
arXiv:1605.01233
Abstract
We formulate an info-clustering paradigm based on a multivariate information measure, called multivariate mutual information, that naturally extends Shannon's mutual information between two random variables to the multivariate case involving more than two random variables. With proper model reductions, we show that the paradigm can be applied to study the human genome and connectome in a more meaningful way than the conventional algorithmic approach. Not only can info-clustering provide justifications and refinements to some existing techniques, but it also inspires new computationally feasible solutions.
In celebration of Claude Shannon's Centenary
References in corpus (8)
- Maps of random walks on complex networks reveal community structure
- Hierarchical Clustering Based on Mutual Information
- Discovering Structure in High-Dimensional Data Through Correlation Explanation
- Parallel Correlation Clustering on Big Graphs
- On the Public Communication Needed to Achieve SK Capacity in the Multiterminal Source Model
- On Tightness of Mutual Dependence Upperbound for Secret-key Capacity of Multiple Terminals
- Achieving SK Capacity in the Source Model: When Must All Terminals Talk?
- Duality between Feature Selection and Data Clustering