Understanding intermediate layers using linear classifier probes
arXiv:1610.01644
Abstract
Neural network models have a reputation for being black boxes. We propose to monitor the features at every layer of a model and measure how suitable they are for classification. We use linear classifiers, which we refer to as "probes", trained entirely independently of the model itself. This helps us better understand the roles and dynamics of the intermediate layers. We demonstrate how this can be used to develop a better intuition about models and to diagnose potential problems. We apply this technique to the popular models Inception v3 and Resnet-50. Among other things, we observe experimentally that the linear separability of features increase monotonically along the depth of the model.
Cited by in corpus (32)
- On the Opportunities and Risks of Foundation Models
- Explainable Artificial Intelligence: a Systematic Review
- Open-Ended Learning Leads to Generally Capable Agents
- The emergence of number and syntax units in LSTM language models
- Boosting Occluded Image Classification via Subspace Decomposition Based Estimation of Deep Features
- Pareto Probing: Trading Off Accuracy for Complexity
- Unsupervised State Representation Learning in Atari
- AI-AI Bias: large language models favor communications generated by large language models
- Deep Learning Through the Lens of Example Difficulty
- Designing and Interpreting Probes with Control Tasks
- Deep Convolutional Decision Jungle for Image Classification
- Evaluating representations by the complexity of learning low-loss predictors
- On the importance of cross-task features for class-incremental learning
- Action-Based Representation Learning for Autonomous Driving
- Screening Gender Transfer in Neural Machine Translation
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- Internal representation dynamics and geometry in recurrent neural networks
- Example-Based Concept Analysis Framework for Deep Weather Forecast Models
- An information theoretic view on selecting linguistic probes
- DirectProbe: Studying Representations without Classifiers
- What BERT Based Language Models Learn in Spoken Transcripts: An Empirical Study
- Prior Activation Distribution (PAD): A Versatile Representation to Utilize DNN Hidden Units
- Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
- Do Syntactic Probes Probe Syntax? Experiments with Jabberwocky Probing
- Explainability-aided Domain Generalization for Image Classification
- Cause and Effect: Hierarchical Concept-based Explanation of Neural Networks
- Scrutinizing and De-Biasing Intuitive Physics with Neural Stethoscopes
- Logic and the -Simplicial Transformer
- ReLU Code Space: A Basis for Rating Network Quality Besides Accuracy
- Conditional probing: measuring usable information beyond a baseline
- Examining the rhetorical capacities of neural language models