Algorithms for Learning Kernels Based on Centered Alignment
arXiv:1203.0550
Abstract
This paper presents new and effective algorithms for learning kernels. In particular, as shown by our empirical results, these algorithms consistently outperform the so-called uniform combination solution that has proven to be difficult to improve upon in the past, as well as other algorithms for learning kernels based on convex combinations of base kernels in both classification and regression. Our algorithms are based on the notion of centered alignment which is used as a similarity measure between kernels or kernel matrices. We present a number of novel algorithmic, theoretical, and empirical results for learning kernels based on our notion of centered alignment. In particular, we describe efficient algorithms for learning a maximum alignment kernel by showing that the problem can be reduced to a simple QP and discuss a one-stage algorithm for learning both a kernel and a hypothesis based on that kernel using an alignment-based regularization. Our theoretical results include a novel concentration bound for centered alignment between kernel matrices, the proof of the existence of effective predictors for kernels with high alignment, both for classification and for regression, and the proof of stability-based generalization bounds for a broad family of algorithms for learning kernels based on centered alignment. We also report the results of experiments with our centered alignment-based algorithms in both classification and regression.
References in corpus (2)
Cited by in corpus (59)
- Training Quantum Embedding Kernels on Near-Term Quantum Computers
- Do Vision Transformers See Like Convolutional Neural Networks?
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- Large-Scale Kernel Methods for Independence Testing
- Entangled Watermarks as a Defense against Model Extraction
- Graph Convolutional Network-based Feature Selection for High-dimensional and Low-sample Size Data
- Learning Attributes Equals Multi-Source Domain Generalization
- The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
- Supervised LogEuclidean Metric Learning for Symmetric Positive Definite Matrices
- Similarity of Neural Network Models: A Survey of Functional and Representational Measures
- Self-Supervised Learning with Kernel Dependence Maximization
- Gaussian RBF Centered Kernel Alignment (CKA) in the Large Bandwidth Limit
- Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman Kernels
- Quantum Kernel Alignment with Stochastic Gradient Descent
- Similarity of Neural Networks with Gradients
- Scale-variant topological information for characterizing the structure of complex networks
- Multi-View Spectral Clustering with High-Order Optimal Neighborhood Laplacian Matrix
- Ultra High-Dimensional Nonlinear Feature Selection for Big Biological Data
- An Empirical Study on Post-processing Methods for Word Embeddings
- Learning the kernel matrix via predictive low-rank approximations
- Topological Persistence Guided Knowledge Distillation for Wearable Sensor Data
- Teachers Do More Than Teach: Compressing Image-to-Image Models
- Risk Convergence of Centered Kernel Ridge Regression with Large Dimensional Data
- ExpPoint-MAE: Better interpretability and performance for self-supervised point cloud transformers
- Implicit Regularization via Neural Feature Alignment
- The Inductive Bias of Quantum Kernels
- Benchmarking Spiking Neural Network Learning Methods with Varying Locality
- Multi-view Unsupervised Feature Selection by Cross-diffused Matrix Alignment
- Optimization and Generalization Analysis of Transduction through Gradient Boosting and Application to Multi-scale Graph Neural Networks
- PIVOT- Input-aware Path Selection for Energy-efficient ViT Inference
- A low variance consistent test of relative dependency
- Solving Interpretable Kernel Dimension Reduction
- Kernel-based parameter estimation of dynamical systems with unknown observation functions
- A measure of association between vectors based on "similarity covariance"
- Analysis of Knowledge Transfer in Kernel Regime
- NLARS: Minimum Redundancy Maximum Relevance Feature Selection for Large and High-dimensional Data
- Graph-Based Similarity of Neural Network Representations
- Deep Transfer Learning with Ridge Regression
- Latent regularization for feature selection using kernel methods in tumor classification
- Structured Prediction by Conditional Risk Minimization
- On the Expressive Power of Kernel Methods and the Efficiency of Kernel Learning by Association Schemes
- Geometry-aware Similarity Learning on SPD Manifolds for Visual Recognition
- Why Do Better Loss Functions Lead to Less Transferable Features?
- A Distributionally Robust Optimization Method for Adversarial Multiple Kernel Learning
- Principled Non-Linear Feature Selection
- Learning Non-Linear Feature Maps
- Neural Networks as Kernel Learners: The Silent Alignment Effect
- Multi-view Kernel Completion
- Learning Fair Canonical Polyadical Decompositions using a Kernel Independence Criterion
- Random Fourier Features via Fast Surrogate Leverage Weighted Sampling
- Exploiting a Zoo of Checkpoints for Unseen Tasks
- Kernel Dependence Network
- Entangled Kernels -- Beyond Separability
- Positive semidefinite support vector regression metric learning
- Correlations between Word Vector Sets
- Meta-Learning to Improve Pre-Training
- Generalization Properties of hyper-RKHS and its Applications
- Learning Data-adaptive Nonparametric Kernels
- Optimality Implies Kernel Sum Classifiers are Statistically Efficient