Taking Advantage of Sparsity in Multi-Task Learning
arXiv:0903.1468
Abstract
We study the problem of estimating multiple linear regression equations for the purpose of both prediction and variable selection. Following recent work on multi-task learning Argyriou et al. [2008], we assume that the regression vectors share the same sparsity pattern. This means that the set of relevant predictor variables is the same across the different equations. This assumption leads us to consider the Group Lasso as a candidate estimation method. We show that this estimator enjoys nice sparsity oracle inequalities and variable selection properties. The results hold under a certain restricted eigenvalue condition and a coherence condition on the design matrix, which naturally extend recent work in Bickel et al. [2007], Lounici [2008]. In particular, in the multi-task learning scenario, in which the number of tasks can grow, we are able to remove completely the effect of the number of predictor variables in the bounds. Finally, we show how our results can be extended to more general noise distributions, of which we only require the variance to be finite.
References in corpus (1)
Cited by in corpus (32)
- Least squares after model selection in high-dimensional sparse models
- Square-Root Lasso: Pivotal Recovery of Sparse Signals via Conic Programming
- Multi-Task Learning with Deep Neural Networks: A Survey
- Learning with Structured Sparsity
- Regularization Techniques for Learning with Matrices
- Learn on Source, Refine on Target:A Model Transfer Learning Framework with Random Forests
- Pivotal estimation via square-root Lasso in nonparametric regression
- Multi-Stage Multi-Task Feature Learning
- L1-Penalized Quantile Regression in High-Dimensional Sparse Models
- Learning the Conditional Independence Structure of Stationary Time Series: A Multitask Learning Approach
- Inferring large graphs using l1-penalized likelihood
- When Deep Learning Meets Multi-Task Learning in SAR ATR: Simultaneous Target Recognition and Segmentation
- Scalable Transfer Learning with Expert Models
- Exact block-wise optimization in group lasso and sparse group lasso for linear regression
- Transfer Learning for High-dimensional Linear Regression: Prediction, Estimation, and Minimax Optimality
- Asymptotic Analysis of Complex LASSO via Complex Approximate Message Passing (CAMP)
- A General Framework of Dual Certificate Analysis for Structured Sparse Recovery Problems
- Structured Sparse Regression via Greedy Hard-Thresholding
- Fast global convergence of gradient methods for high-dimensional statistical recovery
- Maximin Analysis of Message Passing Algorithms for Recovering Block Sparse Signals
- Classification with Sparse Overlapping Groups
- Generalized Kalman Smoothing: Modeling and Algorithms
- Transfer Learning for Linear Regression: a Statistical Test of Gain
- Collaborative Multi-sensor Classification via Sparsity-based Representation
- Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
- Support recovery and sup-norm convergence rates for sparse pivotal estimation
- Transfer Learning in Information Criteria-based Feature Selection
- Simultaneous support recovery in high dimensions: Benefits and perils of block -regularization
- Context-dependent self-exciting point processes: models, methods, and risk bounds in high dimensions
- Bounds of restricted isometry constants in extreme asymptotics: formulae for Gaussian matrices
- Sparse Empirical Bayes Analysis (SEBA)
- Exploring Sparsity in Multi-class Linear Discriminant Analysis