Generalization Properties of Learning with Random Features
arXiv:1602.04474
Abstract
We study the generalization properties of ridge regression with random features in the statistical learning framework. We show for the first time that learning bounds can be achieved with only random features rather than as suggested by previous results. Further, we prove faster learning rates and show that they might require more random features, unless they are sampled according to a possibly problem dependent distribution. Our results shed light on the statistical computational trade-offs in large scale kernelized learning, showing the potential effectiveness of random features in reducing the computational complexity while keeping optimal generalization properties.
NIPS 2017
Cited by in corpus (85)
- Reconciling modern machine learning practice and the bias-variance trade-off
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Domain Generalization by Marginal Transfer Learning
- Probabilistic Load Forecasting Based on Adaptive Online Learning
- Batched Large-scale Bayesian Optimization in High-dimensional Spaces
- Optimal Rates for Spectral Algorithms with Least-Squares Regression over Hilbert Spaces
- Are we done with object recognition? The iCub robot's perspective
- Generalization Bounds for Sparse Random Feature Expansions
- Multiple Descent: Design Your Own Generalization Curve
- Kernel methods through the roof: handling billions of points efficiently
- Learning with invariances in random features and kernel models
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- A Random Matrix Analysis of Random Fourier Features: Beyond the Gaussian Kernel, a Precise Phase Transition, and the Corresponding Double Descent
- Generalization error of random features and kernel methods: hypercontractivity and kernel matrix concentration
- On the Approximation Properties of Random ReLU Features
- Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel Regime
- On Fast Leverage Score Sampling and Optimal Learning
- Approximate Kernel PCA Using Random Features: Computational vs. Statistical Trade-off
- Kernel regression in high dimensions: Refined analysis beyond double descent
- Reservoir Computing meets Recurrent Kernels and Structured Transforms
- Limitations of Lazy Training of Two-layers Neural Networks
- When Does Preconditioning Help or Hurt Generalization?
- Random Feature-based Online Multi-kernel Learning in Environments with Unknown Dynamics
- Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
- Learning Strategies in Decentralized Matching Markets under Uncertain Preferences
- Invariance of Weight Distributions in Rectified MLPs
- Deep Equals Shallow for ReLU Networks in Kernel Regimes
- Implicit Regularization of Random Feature Models
- On the Estimation of Derivatives Using Plug-in Kernel Ridge Regression Estimators
- Risk Convergence of Centered Kernel Ridge Regression with Large Dimensional Data
- A General Framework for Consistent Structured Prediction with Implicit Loss Embeddings
- COKE: Communication-Censored Decentralized Kernel Learning
- Conditioning of Random Feature Matrices: Double Descent and Generalization Error
- Random feature neural networks learn Black-Scholes type PDEs without curse of dimensionality
- On Sampling Random Features From Empirical Leverage Scores: Implementation and Theoretical Guarantees
- Benign Overfitting and Noisy Features
- Breaking the waves: asymmetric random periodic features for low-bitrate kernel machines
- Statistical Estimation of the Poincar{é} constant and Application to Sampling Multimodal Distributions
- On Kernel Derivative Approximation with Random Fourier Features
- On the Double Descent of Random Features Models Trained with SGD
- PSD Representations for Effective Probability Models
- Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models
- Towards A Unified Analysis of Random Fourier Features
- Simple and Almost Assumption-Free Out-of-Sample Bound for Random Feature Mapping
- McKernel: A Library for Approximate Kernel Expansions in Log-linear Time
- Data-driven Random Fourier Features using Stein Effect
- Improved Classification Rates for Localized SVMs
- Efficient online learning with kernels for adversarial large scale problems
- Scalable Global Alignment Graph Kernel Using Random Features: From Node Embedding to Graph Embedding
- Gradient Descent in RKHS with Importance Labeling
- Relating Leverage Scores and Density using Regularized Christoffel Functions
- Generalization Guarantees for Sparse Kernel Approximation with Entropic Optimal Features
- Graph Random Neural Features for Distance-Preserving Graph Representations
- The Error Probability of Random Fourier Features is Dimensionality Independent
- Towards Sharp Analysis for Distributed Learning with Random Features
- Scaling Neural Tangent Kernels via Sketching and Random Features
- A Conceptual Framework for Lifelong Learning
- Gain with no Pain: Efficient Kernel-PCA by Nyström Sampling
- Exponential Error Convergence in Data Classification with Optimized Random Features: Acceleration by Quantum Machine Learning
- Pseudo-Bayesian Learning with Kernel Fourier Transform as Prior
- Random Network Distillation as a Diversity Metric for Both Image and Text Generation
- Sigma-Delta and Distributed Noise-Shaping Quantization Methods for Random Fourier Features
- Nearly Optimal Clustering Risk Bounds for Kernel K-Means
- Exponential Convergence Rates of Classification Errors on Learning with SGD and Random Features
- ORCCA: Optimal Randomized Canonical Correlation Analysis
- Additive function approximation in the brain
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Learning Augmentation Distributions using Transformed Risk Minimization
- Statistical Optimality and Computational Efficiency of Nyström Kernel PCA
- Sampling from Arbitrary Functions via PSD Models
- Convolutional Spectral Kernel Learning
- Over-parametrized neural networks as under-determined linear systems
- Generalization Properties of hyper-RKHS and its Applications
- Denoising Score Matching with Random Fourier Features
- Supervising Nyström Methods via Negative Margin Support Vector Selection
- Regularized ERM on random subspaces
- Random Features for the Neural Tangent Kernel
- Cell Association via Boundary Detection: A Scalable Approach Based on Data-Driven Random Features
- Random Fourier Features via Fast Surrogate Leverage Weighted Sampling
- Efficient Global String Kernel with Random Features: Beyond Counting Substructures
- The Separation Capacity of Random Neural Networks
- Large-scale Kernel Methods and Applications to Lifelong Robot Learning
- ParK: Sound and Efficient Kernel Ridge Regression by Feature Space Partitions
- Shallow Representation is Deep: Learning Uncertainty-aware and Worst-case Random Feature Dynamics
- Minimum complexity interpolation in random features models