Optimal learning with -aggregation
arXiv:1301.6080 · doi:10.1214/13-AOS1190
Abstract
We consider a general supervised learning problem with strongly convex and Lipschitz loss and study the problem of model selection aggregation. In particular, given a finite dictionary functions (learners) together with the prior, we generalize the results obtained by Dai, Rigollet and Zhang [Ann. Statist. 40 (2012) 1878-1905] for Gaussian regression with squared loss and fixed design to this learning setup. Specifically, we prove that the -aggregation procedure outputs an estimator that satisfies optimal oracle inequalities both in expectation and with high probability. Our proof techniques somewhat depart from traditional proofs by making most of the standard arguments on the Laplace transform of the empirical process to be controlled.
Published in at http://dx.doi.org/10.1214/13-AOS1190 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (7)
- Aggregation for Gaussian regression
- Fast learning rates in statistical inference through aggregation
- Sparse Estimation by Exponential Weighting
- Deviation optimal learning using greedy Q-aggregation
- Pac-bayesian bounds for sparse regression estimation with exponential weights
- Optimal rates of aggregation in classification under low noise assumption
- Sharper lower bounds on the performance of the empirical risk minimization algorithm
Cited by in corpus (6)
- Empirical entropy, minimax regret and minimax risk
- Performance of empirical risk minimization in linear aggregation
- Optimal bounds for aggregation of affine estimators
- Robust angle-based transfer learning in high dimensions
- Optimal exponential bounds for aggregation of density estimators
- An adaptive multiclass nearest neighbor classifier