Similarity of Neural Network Representations Revisited
arXiv:1905.00414
Abstract
Recent work has sought to understand the behavior of neural networks by comparing representations between layers and between different trained models. We examine methods for comparing neural network representations based on canonical correlation analysis (CCA). We show that CCA belongs to a family of statistics for measuring multivariate similarity, but that neither CCA nor any other statistic that is invariant to invertible linear transformation can measure meaningful similarities between representations of higher dimension than the number of data points. We introduce a similarity index that measures the relationship between representational similarity matrices and does not suffer from this limitation. This similarity index is equivalent to centered kernel alignment (CKA) and is also closely connected to CCA. Unlike CCA, CKA can reliably identify correspondences between representations in networks trained from different initializations.
ICML 2019
Cited by in corpus (107)
- Convolutional Neural Networks as a Model of the Visual System: Past, Present, and Future
- Artificial neural networks for neuroscientists: A primer
- What is being transferred in transfer learning?
- Do Vision Transformers See Like Convolutional Neural Networks?
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and Depth
- On Representation Knowledge Distillation for Graph Neural Networks
- BOIL: Towards Representation Change for Few-shot Learning
- A Survey of Deep Learning for Scientific Discovery
- What shapes feature representations? Exploring datasets, architectures, and training
- Adaptive Model Pooling for Online Deep Anomaly Detection from a Complex Evolving Data Stream
- Self-Supervised Pre-Training for Transformer-Based Person Re-Identification
- Learning Student-Friendly Teacher Networks for Knowledge Distillation
- Entangled Watermarks as a Defense against Model Extraction
- Advancing diagnostic performance and clinical usability of neural networks via adversarial training and dual batch normalization
- No Fear of Heterogeneity: Classifier Calibration for Federated Learning with Non-IID Data
- Universal Paralinguistic Speech Representations Using Self-Supervised Conformers
- Quantifying the LiDAR Sim-to-Real Domain Shift: A Detailed Investigation Using Object Detectors and Analyzing Point Clouds at Target-Level
- LGViT: Dynamic Early Exiting for Accelerating Vision Transformer
- Low-Dimensional Structure in the Space of Language Representations is Reflected in Brain Responses
- Supervised Transfer Learning at Scale for Medical Imaging
- Revisiting Model Stitching to Compare Neural Representations
- Radio2Text: Streaming Speech Recognition Using mmWave Radio Signals
- Unraveling Meta-Learning: Understanding Feature Representations for Few-Shot Tasks
- Object Detector Differences when using Synthetic and Real Training Data
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Similarity of Neural Networks with Gradients
- The effect of task and training on intermediate representations in convolutional neural networks revealed with modified RV similarity analysis
- Hierarchically Organized Latent Modules for Exploratory Search in Morphogenetic Systems
- The intriguing role of module criticality in the generalization of deep networks
- The Role of Complex NLP in Transformers for Text Ranking?
- Linear Mode Connectivity in Multitask and Continual Learning
- An Information Theory-inspired Strategy for Automatic Network Pruning
- Machine Learning for Cataract Classification and Grading on Ophthalmic Imaging Modalities: A Survey
- Context Meta-Reinforcement Learning via Neuromodulation
- Is Supervised Syntactic Parsing Beneficial for Language Understanding? An Empirical Investigation
- OMPQ: Orthogonal Mixed Precision Quantization
- Taxonomizing local versus global structure in neural network loss landscapes
- TRS: Transferability Reduced Ensemble via Encouraging Gradient Diversity and Model Smoothness
- PURSUhInT: In Search of Informative Hint Points Based on Layer Clustering for Knowledge Distillation
- Teachers Do More Than Teach: Compressing Image-to-Image Models
- Insights on Neural Representations for End-to-End Speech Recognition
- On Negative Interference in Multilingual Models: Findings and A Meta-Learning Treatment
- VendorLink: An NLP approach for Identifying & Linking Vendor Migrants & Potential Aliases on Darknet Markets
- Transferred Discrepancy: Quantifying the Difference Between Representations
- Nondeterminism and Instability in Neural Network Optimization
- Exploring the Interchangeability of CNN Embedding Spaces
- Harmonizing Feature Attributions Across Deep Learning Architectures: Enhancing Interpretability and Consistency
- SoK: How Robust is Image Classification Deep Neural Network Watermarking? (Extended Version)
- Simon Says: Evaluating and Mitigating Bias in Pruned Neural Networks with Knowledge Distillation
- Grounding Representation Similarity with Statistical Testing
- Exploring Linguistic Properties of Monolingual BERTs with Typological Classification among Languages
- A Channel Coding Benchmark for Meta-Learning
- Cross-Layer Distillation with Semantic Calibration
- Robust Distance Covariance
- Compare Where It Matters: Using Layer-Wise Regularization To Improve Federated Learning on Heterogeneous Data
- FEAR: A Simple Lightweight Method to Rank Architectures
- Graph-Based Similarity of Neural Network Representations
- Undivided Attention: Are Intermediate Layers Necessary for BERT?
- Invariance Measures for Neural Networks
- When to Prune? A Policy towards Early Structural Pruning
- ACDC: Weight Sharing in Atom-Coefficient Decomposed Convolution
- Representation Transfer by Optimal Transport
- Phone and speaker spatial organization in self-supervised speech representations
- NomMer: Nominate Synergistic Context in Vision Transformer for Visual Recognition
- Do Self-Supervised and Supervised Methods Learn Similar Visual Representations?
- Knowledge Distillation Meets Self-Supervision
- Hierarchical nucleation in deep neural networks
- Net2Brain: A Toolbox to compare artificial vision models with human brain responses
- Aggregative Self-Supervised Feature Learning from a Limited Sample
- Revisiting Hidden Representations in Transfer Learning for Medical Imaging
- Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)
- Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
- Interpreting and Disentangling Feature Components of Various Complexity from DNNs
- Exploring the limits of pre-trained embeddings in machine-guided protein design: a case study on predicting AAV vector viability
- Identifying the Correlation Between Language Distance and Cross-Lingual Transfer in a Multilingual Representation Space
- Trivial or impossible -- dichotomous data difficulty masks model differences (on ImageNet and beyond)
- Semi-Online Knowledge Distillation
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
- Comparison Against Task Driven Artificial Neural Networks Reveals Functional Organization of Mouse Visual Cortex
- ImageNet Pre-training also Transfers Non-Robustness
- Properties of the After Kernel
- How do Quadratic Regularizers Prevent Catastrophic Forgetting: The Role of Interpolation
- FedProf: Selective Federated Learning with Representation Profiling
- Multi-level Knowledge Distillation via Knowledge Alignment and Correlation
- Why Do Better Loss Functions Lead to Less Transferable Features?
- Similarity Analysis of Self-Supervised Speech Representations
- Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual Interpolation
- Improved Knowledge Distillation via Adversarial Collaboration
- Meta-Learning to Improve Pre-Training
- LOOPER: Inferring computational algorithms enacted by neuronal population dynamics
- Exploiting a Zoo of Checkpoints for Unseen Tasks
- How Familiar Does That Sound? Cross-Lingual Representational Similarity Analysis of Acoustic Word Embeddings
- Correlations between Word Vector Sets
- Understanding the Dynamics of DNNs Using Graph Modularity
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- Multi-granularity for knowledge distillation
- A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling
- One-Shot Neural Ensemble Architecture Search by Diversity-Guided Search Space Shrinking
- Disentangling Transfer and Interference in Multi-Domain Learning
- An Investigation of the Weight Space to Monitor the Training Progress of Neural Networks
- Pull-back Geometry of Persistent Homology Encodings
- On Cross-Layer Alignment for Model Fusion of Heterogeneous Neural Networks
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models
- Reducing the Human Effort in Developing PET-CT Registration
- OrthoReg: Robust Network Pruning Using Orthonormality Regularization
- Characterizing and Measuring the Similarity of Neural Networks with Persistent Homology
- Does the Adam Optimizer Exacerbate Catastrophic Forgetting?