Learning with SGD and Random Features
arXiv:1807.06343
Abstract
Sketching and stochastic gradient methods are arguably the most common techniques to derive efficient large scale learning algorithms. In this paper, we investigate their application in the context of nonparametric statistical learning. More precisely, we study the estimator defined by stochastic gradient with mini batches and random features. The latter can be seen as form of nonlinear sketching and used to define approximate kernel methods. The considered estimator is not explicitly penalized/constrained and regularization is implicit. Indeed, our study highlights how different parameters, such as number of features, iterations, step-size and mini-batch size control the learning properties of the solutions. We do this by deriving optimal finite sample bounds, under standard assumptions. The obtained results are corroborated and illustrated by numerical experiments.
Cited by in corpus (20)
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Analysis of the Gradient Descent Algorithm for a Deep Neural Network Model with Skip-connections
- On the Approximation Properties of Random ReLU Features
- Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime
- On Fast Leverage Score Sampling and Optimal Learning
- Reservoir Computing meets Recurrent Kernels and Structured Transforms
- Optimal Statistical Rates for Decentralised Non-Parametric Regression with Linear Speed-Up
- A General Framework for Consistent Structured Prediction with Implicit Loss Embeddings
- Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality
- On Sampling Random Features From Empirical Leverage Scores: Implementation and Theoretical Guarantees
- The Slow Deterioration of the Generalization Error of the Random Feature Model
- On the Double Descent of Random Features Models Trained with SGD
- Gradient Descent in RKHS with Importance Labeling
- Improved Classification Rates for Localized SVMs
- Exponential Error Convergence in Data Classification with Optimized Random Features: Acceleration by Quantum Machine Learning
- Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features
- Dimension Independent Generalization Error by Stochastic Gradient Descent
- Generalization Properties of hyper-RKHS and its Applications
- ParK: Sound and Efficient Kernel Ridge Regression by Feature Space Partitions
- Shallow Representation is Deep: Learning Uncertainty-aware and Worst-case Random Feature Dynamics