Optimal learning via local entropies and sample compression
arXiv:1706.01124
Abstract
The aim of this paper is to provide several novel upper bounds on the excess risk with a primal focus on classification problems. We suggest two approaches and the obtained bounds are represented via the distribution dependent local entropies of the classes or the sizes of specific sample com- pression schemes. We show that in some cases, our guarantees are optimal up to constant factors and outperform previously known results. As an application of our results, we provide a new tight PAC bound for the hard-margin SVM, an extended analysis of certain empirical risk minimizers under log-concave distributions, a new variant of an online to batch conversion, and distribution dependent localized bounds in the aggregation framework. We also develop techniques that allow to replace empirical covering number or covering numbers with bracketing by the coverings with respect to the distribution of the data. The proofs for the sample compression schemes are based on the moment method combined with the analysis of voting algorithms.
25 pages. Extended and restructured version. Contains new results, reflected in the new section 4 and section 5. Corrects Lemma 4 of the previous version. Sample compression part remained unchanged
References in corpus (3)
Cited by in corpus (6)
- Sharper bounds for uniformly stable algorithms
- Proper Learning, Helly Number, and an Optimal SVM Bound
- Stable Sample Compression Schemes: New Applications and an Optimal SVM Margin Bound
- Towards Optimal Problem Dependent Generalization Error Bounds in Statistical Learning Theory
- When are epsilon-nets small?
- A New Lower Bound for Agnostic Learning with Sample Compression Schemes