Profile Entropy: A Fundamental Measure for the Learnability and Compressibility of Discrete Distributions
arXiv:2002.11665
Abstract
The profile of a sample is the multiset of its symbol frequencies. We show that for samples of discrete distributions, profile entropy is a fundamental measure unifying the concepts of estimation, inference, and compression. Specifically, profile entropy a) determines the speed of estimating the distribution relative to the best natural estimator; b) characterizes the rate of inferring all symmetric properties compared with the best estimator over any label-invariant distribution collection; c) serves as the limit of profile compression, for which we derive optimal near-linear-time block and sequential algorithms. To further our understanding of profile entropy, we investigate its attributes, provide algorithms for approximating its value, and determine its magnitude for numerous structural distribution families.
56 pages
References in corpus (7)
- Entropy inference and the James-Stein estimator, with application to nonlinear gene association networks
- Monotone probability distributions over the Boolean cube can be learned with sublinear samples
- Differentially private anonymized histograms
- Data Amplification: A Unified and Competitive Approach to Property Estimation
- Data Amplification: Instance-Optimal Property Estimation
- Sublinear Optimal Policy Value Estimation in Contextual Bandits
- Phase Transitions for the Uniform Distribution in the PML Problem and its Bethe Approximation