Recycling Randomness with Structure for Sublinear time Kernel Expansions
arXiv:1605.09049
Abstract
We propose a scheme for recycling Gaussian random vectors into structured matrices to approximate various kernel functions in sublinear time via random embeddings. Our framework includes the Fastfood construction as a special case, but also extends to Circulant, Toeplitz and Hankel matrices, and the broader family of structured matrices that are characterized by the concept of low-displacement rank. We introduce notions of coherence and graph-theoretic structural constants that control the approximation quality, and prove unbiasedness and low-variance properties of random feature maps that arise within our framework. For the case of low-displacement matrices, we show how the degree of structure and randomness can be controlled to reduce statistical variance at the cost of increased computation and storage requirements. Empirical results strongly support our theory and justify the use of a broader family of structured matrices for scaling up kernel methods using random features.
References in corpus (5)
Cited by in corpus (12)
- Structured adaptive and random spinners for fast machine learning computations
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Local Group Invariant Representations via Orbit Embeddings
- Sketching for Large-Scale Learning of Mixture Models
- On the Expressive Power of Self-Attention Matrices
- Learning Compressed Transforms with Low Displacement Rank
- Matrix Infinitely Divisible Series: Tail Inequalities and Their Applications
- On Learning the Transformer Kernel
- Revisiting RIP guarantees for sketching operators on mixture models
- TripleSpin - a generic compact paradigm for fast machine learning computations
- Demystifying Orthogonal Monte Carlo and Beyond
- On Dimension-free Tail Inequalities for Sums of Random Matrices and Applications