Fast learning rate of multiple kernel learning: Trade-off between sparsity and smoothness
arXiv:1203.0565 · doi:10.1214/13-AOS1095
Abstract
We investigate the learning rate of multiple kernel learning (MKL) with and elastic-net regularizations. The elastic-net regularization is a composition of an -regularizer for inducing the sparsity and an -regularizer for controlling the smoothness. We focus on a sparse setting where the total number of kernels is large, but the number of nonzero components of the ground truth is relatively small, and show sharper convergence rates than the learning rates have ever shown for both and elastic-net regularizations. Our analysis reveals some relations between the choice of a regularization function and the performance. If the ground truth is smooth, we show a faster convergence rate for the elastic-net regularization with less conditions than -regularization; otherwise, a faster convergence rate for the -regularization is shown.
Published in at http://dx.doi.org/10.1214/13-AOS1095 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org). arXiv admin note: text overlap with arXiv:1103.0431
References in corpus (7)
- Consistency of the group Lasso and multiple kernel learning
- L2 Regularization for Learning Kernels
- Sparsity in multiple kernel learning
- Exploring Large Feature Spaces with Hierarchical Multiple Kernel Learning
- Fast learning rate of multiple kernel learning: Trade-off between sparsity and smoothness
- A Unifying View of Multiple Kernel Learning
- Fast Learning Rate of Non-Sparse Multiple Kernel Learning and Optimal Regularization Strategies
Cited by in corpus (9)
- High-Dimensional Feature Selection by Feature-Wise Kernelized Lasso
- Fast learning rate of multiple kernel learning: Trade-off between sparsity and smoothness
- PAC-Bayesian Estimation and Prediction in Sparse Additive Models
- Doubly Decomposing Nonparametric Tensor Regression
- Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data
- Minimax Optimal Rates of Estimation in High Dimensional Additive Models: Universal Phase Transition
- Minimax Optimal Estimation in Partially Linear Additive Models under High Dimension
- Extreme Eigenvalues of Nonlinear Correlation Matrices with Applications to Additive Models
- Learning rates for the risk of kernel based quantile regression estimators in additive models