Binary embeddings with structured hashed projections
arXiv:1511.05212
Abstract
We consider the hashing mechanism for constructing binary embeddings, that involves pseudo-random projections followed by nonlinear (sign function) mappings. The pseudo-random projection is described by a matrix, where not all entries are independent random variables but instead a fixed "budget of randomness" is distributed across the matrix. Such matrices can be efficiently stored in sub-quadratic or even linear space, provide reduction in randomness usage (i.e. number of required random values), and very often lead to computational speed ups. We prove several theoretical results showing that projections via various structured matrices followed by nonlinear mappings accurately preserve the angular distance between input high-dimensional vectors. To the best of our knowledge, these results are the first that give theoretical ground for the use of general structured matrices in the nonlinear setting. In particular, they generalize previous extensions of the Johnson-Lindenstrauss lemma and prove the plausibility of the approach that was so far only heuristically confirmed for some special structured matrices. Consequently, we show that many structured matrices can be used as an efficient information compression mechanism. Our findings build a better understanding of certain deep architectures, which contain randomly weighted and untrained layers, and yet achieve high performance on different learning tasks. We empirically verify our theoretical findings and show the dependence of learning via structured hashed projections on the performance of neural network as well as nearest neighbor classifier.
arXiv admin note: text overlap with arXiv:1505.03190
References in corpus (7)
- The Loss Surfaces of Multilayer Networks
- Compressing Neural Networks with the Hashing Trick
- Experiments with Random Projection
- Deep Neural Networks with Random Gaussian Weights: A Universal Classification Strategy?
- Circulant Binary Embedding
- Structured Transforms for Small-Footprint Deep Learning
- Binary Embedding: Fundamental Limits and Fast Algorithm
Cited by in corpus (13)
- The Unreasonable Effectiveness of Structured Random Orthogonal Embeddings
- Recycling Randomness with Structure for Sublinear time Kernel Expansions
- Structured adaptive and random spinners for fast machine learning computations
- Provable Filter Pruning for Efficient Neural Networks
- SiPPing Neural Networks: Sensitivity-informed Provable Pruning of Neural Networks
- Making Online Sketching Hashing Even Faster
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- On the Expressive Power of Self-Attention Matrices
- Random Projection in Deep Neural Networks
- Deep Neural Network Approximation using Tensor Sketching
- Faster Binary Embeddings for Preserving Euclidean Distances
- TripleSpin - a generic compact paradigm for fast machine learning computations
- LLC: Accurate, Multi-purpose Learnt Low-dimensional Binary Codes