Understanding Black-box Predictions via Influence Functions
arXiv:1703.04730
Abstract
How can we explain the predictions of a black-box model? In this paper, we use influence functions -- a classic technique from robust statistics -- to trace a model's prediction through the learning algorithm and back to its training data, thereby identifying training points most responsible for a given prediction. To scale up influence functions to modern machine learning settings, we develop a simple, efficient implementation that requires only oracle access to gradients and Hessian-vector products. We show that even on non-convex and non-differentiable models where the theory breaks down, approximations to influence functions can still provide valuable information. On linear models and convolutional neural networks, we demonstrate that influence functions are useful for multiple purposes: understanding model behavior, debugging models, detecting dataset errors, and even creating visually-indistinguishable training-set attacks.
International Conference on Machine Learning, 2017. (This version adds more historical references and fixes typos.)
References in corpus (7)
- TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems
- DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition
- European Union regulations on algorithmic decision-making and a "right to explanation"
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- Poisoning Attacks against Support Vector Machines
- Not Just a Black Box: Learning Important Features Through Propagating Activation Differences
- "Why Should I Trust You?": Explaining the Predictions of Any Classifier
Cited by in corpus (388)
- A Survey on Explainable Artificial Intelligence (XAI): Towards Medical XAI
- On the Opportunities and Risks of Foundation Models
- Interpretable machine learning: definitions, methods, and applications
- Time Series Forecasting With Deep Learning: A Survey
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Explainable artificial intelligence (XAI) in deep learning-based medical image analysis
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- A Survey on the Explainability of Supervised Machine Learning
- Machine learning methods for histopathological image analysis
- Learning to Reweight Examples for Robust Deep Learning
- Machine Learning for Integrating Data in Biology and Medicine: Principles, Practice, and Opportunities
- GNNExplainer: Generating Explanations for Graph Neural Networks
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Towards Explainable Artificial Intelligence
- Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
- A Survey on Data Augmentation for Text Classification
- Analyzing Federated Learning through an Adversarial Lens
- Spectral Signatures in Backdoor Attacks
- One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques
- What Clinicians Want: Contextualizing Explainable Machine Learning for Clinical End Use
- A Review of Challenges and Opportunities in Machine Learning for Health
- Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches
- Towards Out-Of-Distribution Generalization: A Survey
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- Inverting Gradients -- How easy is it to break privacy in federated learning?
- Post-hoc Interpretability for Neural NLP: A Survey
- Mitigating Adversarial Effects Through Randomization
- Fast Yet Effective Machine Unlearning
- Class-Balanced Loss Based on Effective Number of Samples
- Do Adversarially Robust ImageNet Models Transfer Better?
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- Dataset Condensation with Gradient Matching
- Unifying Graph Convolutional Neural Networks and Label Propagation
- TSViz: Demystification of Deep Learning Models for Time-Series Analysis
- Generative Data Augmentation for Commonsense Reasoning
- Hierarchical interpretations for neural network predictions
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation
- MetaPoison: Practical General-purpose Clean-label Data Poisoning
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks
- Long-tailed Visual Recognition via Gaussian Clouded Logit Adjustment
- Demystifying statistical learning based on efficient influence functions
- Explainable Artificial Intelligence: a Systematic Review
- Towards Automatic Concept-based Explanations
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Towards Adversarial Malware Detection: Lessons Learned from PDF-based Attacks
- On the Accuracy of Influence Functions for Measuring Group Effects
- Unsupervised machine learning of topological phase transitions from experimental data
- Counterfactual Explanations for Machine Learning on Multivariate Time Series Data
- Deep Learning for Sequential Recommendation: Algorithms, Influential Factors, and Evaluations
- Keeping Community in the Loop: Understanding Wikipedia Stakeholder Values for Machine Learning-Based Systems
- XAIR: A Framework of Explainable AI in Augmented Reality
- Unnoticeable Backdoor Attacks on Graph Neural Networks
- Interpretation of Neural Networks is Fragile
- Improving Fairness for Data Valuation in Horizontal Federated Learning
- Explainable Machine Learning for Public Policy: Use Cases, Gaps, and Research Directions
- Counterfactual State Explanations for Reinforcement Learning Agents via Generative Deep Learning
- Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Training Data Influence Analysis and Estimation: A Survey
- Federated Unlearning: A Survey on Methods, Design Guidelines, and Evaluation Metrics
- Towards Relatable Explainable AI with the Perceptual Process
- On the (In)fidelity and Sensitivity for Explanations
- EAD: Elastic-Net Attacks to Deep Neural Networks via Adversarial Examples
- Security and Privacy Issues in Deep Learning
- POTs: Protective Optimization Technologies
- GIF: A General Graph Unlearning Strategy via Influence Function
- Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks
- On Completeness-aware Concept-Based Explanations in Deep Neural Networks
- Concise Fuzzy System Modeling Integrating Soft Subspace Clustering and Sparse Learning
- Explaining Latent Factor Models for Recommendation with Influence Functions
- Poisoning Attacks with Generative Adversarial Nets
- On Interpretability of Artificial Neural Networks: A Survey
- Unlearnable Examples: Making Personal Data Unexploitable
- Data Valuation using Reinforcement Learning
- Explain and Predict, and then Predict Again
- Defining Locality for Surrogates in Post-hoc Interpretablity
- Label-Consistent Backdoor Attacks
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- Causality Learning: A New Perspective for Interpretable Machine Learning
- Machine Learning Explainability for External Stakeholders
- Towards a mathematical framework to inform Neural Network modelling via Polynomial Regression
- MAGIX: Model Agnostic Globally Interpretable Explanations
- Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
- Evaluating Explanation Without Ground Truth in Interpretable Machine Learning
- NormLime: A New Feature Importance Metric for Explaining Deep Neural Networks
- Can You Trust This Prediction? Auditing Pointwise Reliability After Learning
- WoodFisher: Efficient Second-Order Approximation for Neural Network Compression
- On Adversarial Bias and the Robustness of Fair Machine Learning
- Adversarial Unlearning of Backdoors via Implicit Hypergradient
- Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
- Can I Trust the Explainer? Verifying Post-hoc Explanatory Methods
- Tree Space Prototypes: Another Look at Making Tree Ensembles Interpretable
- Socially Responsible AI Algorithms: Issues, Purposes, and Challenges
- Machine Learning Security against Data Poisoning: Are We There Yet?
- Adaptive Self-training for Few-shot Neural Sequence Labeling
- Explaining First Impressions: Modeling, Recognizing, and Explaining Apparent Personality from Videos
- Adversarial Examples Make Strong Poisons
- Markov Chain Monte Carlo-Based Machine Unlearning: Unlearning What Needs to be Forgotten
- Coded Machine Unlearning
- Granger-causal Attentive Mixtures of Experts: Learning Important Features with Neural Networks
- A Survey on Neural Network Interpretability
- Multi-modal Deep Guided Filtering for Comprehensible Medical Image Processing
- Factors Influencing Perceived Fairness in Algorithmic Decision-Making: Algorithm Outcomes, Development Procedures, and Individual Differences
- Monitoring and explainability of models in production
- Model-Based Counterfactual Synthesizer for Interpretation
- Explaining and Improving Model Behavior with k Nearest Neighbor Representations
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Adversarial Neuron Pruning Purifies Backdoored Deep Models
- Deep Active Learning for Computer Vision: Past and Future
- IcoRating: A Deep-Learning System for Scam ICO Identification
- Interpreting CNNs via Decision Trees
- Privacy Preservation in Federated Learning: An insightful survey from the GDPR Perspective
- Self-Explaining Structures Improve NLP Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- HYDRA: Hypergradient Data Relevance Analysis for Interpreting Deep Neural Networks
- Beyond Individualized Recourse: Interpretable and Interactive Summaries of Actionable Recourses
- Data Cleansing for Models Trained with SGD
- Explaining Vulnerabilities of Deep Learning to Adversarial Malware Binaries
- Towards Frequency-Based Explanation for Robust CNN
- An explainable deep vision system for animal classification and detection in trail-camera images with automatic post-deployment retraining
- What can AI do for me: Evaluating Machine Learning Interpretations in Cooperative Play
- Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
- Deep Learning on a Data Diet: Finding Important Examples Early in Training
- Dataset Distillation
- Explaining Deep Neural Networks
- Penalty Method for Inversion-Free Deep Bilevel Optimization
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
- Better, Not Just More: Data-Centric Machine Learning for Earth Observation
- Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning
- Certified Data Removal from Machine Learning Models
- Global Model Interpretation via Recursive Partitioning
- Knowledge Consistency between Neural Networks and Beyond
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence Functions
- Poisoning and Backdooring Contrastive Learning
- Certified Robustness to Label-Flipping Attacks via Randomized Smoothing
- A backdoor attack against LSTM-based text classification systems
- Evaluating the Correctness of Explainable AI Algorithms for Classification
- Evaluation of Similarity-based Explanations
- Online Data Poisoning Attack
- Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman Kernels
- Hessian-based toolbox for reliable and interpretable machine learning in physics
- Concentration study of M-estimators using the influence function
- xGEMs: Generating Examplars to Explain Black-Box Models
- Semi-Supervised Learning by Label Gradient Alignment
- Towards Robust, Locally Linear Deep Networks
- Poison as a Cure: Detecting & Neutralizing Variable-Sized Backdoor Attacks in Deep Neural Networks
- ProtoAttend: Attention-Based Prototypical Learning
- Data Dropout: Optimizing Training Data for Convolutional Neural Networks
- Rethinking the Backdoor Attacks' Triggers: A Frequency Perspective
- Generative causal explanations of black-box classifiers
- On Learning and Learned Data Representation by Capsule Networks
- Influence Functions in Deep Learning Are Fragile
- Auditing Data Provenance in Text-Generation Models
- Exact and Consistent Interpretation for Piecewise Linear Neural Networks: A Closed Form Solution
- Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- Decisions, Counterfactual Explanations and Strategic Behavior
- Rewarding High-Quality Data via Influence Functions
- On the Necessity of Auditable Algorithmic Definitions for Machine Unlearning
- Causal Interpretability for Machine Learning -- Problems, Methods and Evaluation
- Explanation by Progressive Exaggeration
- VerIDeep: Verifying Integrity of Deep Neural Networks through Sensitive-Sample Fingerprinting
- JuryGCN: Quantifying Jackknife Uncertainty on Graph Convolutional Networks
- Interpretable Off-Policy Evaluation in Reinforcement Learning by Highlighting Influential Transitions
- Influence-guided Data Augmentation for Neural Tensor Completion
- Toward Better Generalization Bounds with Locally Elastic Stability
- ExplaiNE: An Approach for Explaining Network Embedding-based Link Predictions
- FedCCEA : A Practical Approach of Client Contribution Evaluation for Federated Learning
- Interpretable Machine Learning: Moving From Mythos to Diagnostics
- Personalized explanation in machine learning: A conceptualization
- Rethinking Influence Functions of Neural Networks in the Over-parameterized Regime
- A Higher-Order Swiss Army Infinitesimal Jackknife
- Learning outside the Black-Box: The pursuit of interpretable models
- Understanding Instance-based Interpretability of Variational Auto-Encoders
- PAC-Bayes Information Bottleneck
- Efficient computation of counterfactual explanations of LVQ models
- Certification of embedded systems based on Machine Learning: A survey
- secml: A Python Library for Secure and Explainable Machine Learning
- Morphology of three-body quantum states from machine learning
- Removing biased data to improve fairness and accuracy
- A Marauder's Map of Security and Privacy in Machine Learning
- Sparse Oblique Decision Trees: A Tool to Understand and Manipulate Neural Net Features
- Defending Against Backdoor Attacks in Natural Language Generation
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Variational Prototype Replays for Continual Learning
- The Local Elasticity of Neural Networks
- Chasing Your Long Tails: Differentially Private Prediction in Health Care Settings
- Rationalizing Predictions by Adversarial Information Calibration
- Towards Aggregating Weighted Feature Attributions
- Disentangling Influence: Using Disentangled Representations to Audit Model Predictions
- Algorithmic Recourse in the Wild: Understanding the Impact of Data and Model Shifts
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Regional Tree Regularization for Interpretability in Black Box Models
- Optimizing for Interpretability in Deep Neural Networks with Tree Regularization
- Deep Active Learning by Leveraging Training Dynamics
- Provably efficient, succinct, and precise explanations
- Regularisation Can Mitigate Poisoning Attacks: A Novel Analysis Based on Multiobjective Bilevel Optimisation
- Incorporating Priors with Feature Attribution on Text Classification
- Geometrization of deep networks for the interpretability of deep learning systems
- TracInAD: Measuring Influence for Anomaly Detection
- PANDA: Facilitating Usable AI Development
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
- Model-Targeted Poisoning Attacks with Provable Convergence
- On Data Augmentation and Adversarial Risk: An Empirical Analysis
- FairIF: Boosting Fairness in Deep Learning via Influence Functions with Validation Set Sensitive Attributes
- Towards a Unified Evaluation of Explanation Methods without Ground Truth
- Debugging Tests for Model Explanations
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Robust Attacks against Multiple Classifiers
- Developing Future Human-Centered Smart Cities: Critical Analysis of Smart City Security, Interpretability, and Ethical Challenges
- Bayesian Inference Forgetting
- Mitigating Sybil Attacks on Differential Privacy based Federated Learning
- Visualizing and Understanding Deep Neural Networks in CTR Prediction
- Revealing Perceptible Backdoors, without the Training Set, via the Maximum Achievable Misclassification Fraction Statistic
- Counterfactual Evaluation for Explainable AI
- Cross-Modal Conceptualization in Bottleneck Models
- Learning More From Less: Towards Strengthening Weak Supervision for Ad-Hoc Retrieval
- The Thousand Faces of Explainable AI Along the Machine Learning Life Cycle: Industrial Reality and Current State of Research
- Influence Function based Data Poisoning Attacks to Top-N Recommender Systems
- DeepSeer: Interactive RNN Explanation and Debugging via State Abstraction
- A Backdoor Attack against 3D Point Cloud Classifiers
- Counterfactual Explanations in Sequential Decision Making Under Uncertainty
- Scalable Multitask Learning Using Gradient-based Estimation of Task Affinity
- M2Lens: Visualizing and Explaining Multimodal Models for Sentiment Analysis
- Towards the Unification and Robustness of Perturbation and Gradient Based Explanations
- Towards interpreting ML-based automated malware detection models: a survey
- Mean-Field Approximation to Gaussian-Softmax Integral with Application to Uncertainty Estimation
- High Dimensional Model Explanations: an Axiomatic Approach
- X-SHAP: towards multiplicative explainability of Machine Learning
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- Synthetic-Neuroscore: Using A Neuro-AI Interface for Evaluating Generative Adversarial Networks
- Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder
- Consistent Risk Estimation in Moderately High-Dimensional Linear Regression
- NLS: an accurate and yet easy-to-interpret regression method
- Data Cleaning for Accurate, Fair, and Robust Models: A Big Data - AI Integration Approach
- SSSE: Efficiently Erasing Samples from Trained Machine Learning Models
- Sequential Explanations with Mental Model-Based Policies
- Knowledge-based XAI through CBR: There is more to explanations than models can tell
- A psychophysics approach for quantitative comparison of interpretable computer vision models
- Interpreting search result rankings through intent modeling
- A Unified Framework for Task-Driven Data Quality Management
- Multi-Stage Influence Function
- Learning Effective Representations for Person-Job Fit by Feature Fusion
- Mediation Challenges and Socio-Technical Gaps for Explainable Deep Learning Applications
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- Instance-Based Learning of Span Representations: A Case Study through Named Entity Recognition
- Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation
- Influence Estimation for Generative Adversarial Networks
- DNN2LR: Interpretation-inspired Feature Crossing for Real-world Tabular Data
- Fooling Adversarial Training with Inducing Noise
- Contextual Prediction Difference Analysis for Explaining Individual Image Classifications
- Understanding Adversarial Examples from the Mutual Influence of Images and Perturbations
- DIVINE: Diverse Influential Training Points for Data Visualization and Model Refinement
- Can Information Flows Suggest Targets for Interventions in Neural Circuits?
- Towards Using Data-Influence Methods to Detect Noisy Samples in Source Code Corpora
- A General Framework for Defending Against Backdoor Attacks via Influence Graph
- Regularization Can Help Mitigate Poisoning Attacks... with the Right Hyperparameters
- Data Cleansing for Deep Neural Networks with Storage-efficient Approximation of Influence Functions
- Predicting Language Recovery after Stroke with Convolutional Networks on Stitched MRI
- Interpretable CNNs for Object Classification
- Robust and Explainable Autoencoders for Unsupervised Time Series Outlier Detection---Extended Version
- Architecture Selection via the Trade-off Between Accuracy and Robustness
- Understanding Goal-Oriented Active Learning via Influence Functions
- Channels, Remote Estimation and Queueing Systems With A Utilization-Dependent Component: A Unifying Survey Of Recent Results
- Connecting Interpretability and Robustness in Decision Trees through Separation
- Approximate Cross-Validation for Structured Models
- Leveraging Sparse Linear Layers for Debuggable Deep Networks
- Interpreting Deep Models through the Lens of Data
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging
- Interactive Label Cleaning with Example-based Explanations
- Spectral Roll-off Points Variations: Exploring Useful Information in Feature Maps by Its Variations
- Explaining AlphaGo: Interpreting Contextual Effects in Neural Networks
- Multi-Source Domain Adaptation for Text Classification via DistanceNet-Bandits
- Auditing and Debugging Deep Learning Models via Decision Boundaries: Individual-level and Group-level Analysis
- Annotation-Free Human Sketch Quality Assessment
- EXS: Explainable Search Using Local Model Agnostic Interpretability
- Assessment of the Reliablity of a Model's Decision by Generalizing Attribution to the Wavelet Domain
- Deep Co-Attention Network for Multi-View Subspace Learning
- Consensus-based Interpretable Deep Neural Networks with Application to Mortality Prediction
- Fortifying Toxic Speech Detectors Against Veiled Toxicity
- Enabling SQL-based Training Data Debugging for Federated Learning
- Robust learning under clean-label attack
- A Neuro-AI Interface for Evaluating Generative Adversarial Networks
- On Sample Based Explanation Methods for NLP:Efficiency, Faithfulness, and Semantic Evaluation
- Sampling the "Inverse Set" of a Neuron: An Approach to Understanding Neural Nets
- Going Grayscale: The Road to Understanding and Improving Unlearnable Examples
- SelfExplain: A Self-Explaining Architecture for Neural Text Classifiers
- Estimating informativeness of samples with Smooth Unique Information
- Optimal Piecewise Local-Linear Approximations
- TREX: Tree-Ensemble Representer-Point Explanations
- FineFool: Fine Object Contour Attack via Attention
- Toward the Understanding of Deep Text Matching Models for Information Retrieval
- Interpretable Neural Network Decoupling
- Interpretable Time-series Classification on Few-shot Samples
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Unifying Model Explainability and Robustness via Machine-Checkable Concepts
- Efficient Client Contribution Evaluation for Horizontal Federated Learning
- Explaining the Road Not Taken
- Explanation of Reinforcement Learning Model in Dynamic Multi-Agent System
- Practical Data Poisoning Attack against Next-Item Recommendation
- Dissecting Generation Modes for Abstractive Summarization Models via Ablation and Attribution
- Defending Against Adversarial Denial-of-Service Data Poisoning Attacks
- Learning and Certification under Instance-targeted Poisoning
- EDDA: Explanation-driven Data Augmentation to Improve Explanation Faithfulness
- Interpreting Deep Learning Model Using Rule-based Method
- TRACER: A Framework for Facilitating Accurate and Interpretable Analytics for High Stakes Applications
- DANCE: Enhancing saliency maps using decoys
- ImageNet Pre-training also Transfers Non-Robustness
- Instance-based Deep Transfer Learning
- Using Wavelets to Analyze Similarities in Image-Classification Datasets
- Robust model training and generalisation with Studentising flows
- Deep Neural Networks for Choice Analysis: Architectural Design with Alternative-Specific Utility Functions
- Pareto Navigation Gradient Descent: a First-Order Algorithm for Optimization in Pareto Set
- Quantifying Epistemic Uncertainty in Deep Learning
- A Minimal Intervention Definition of Reverse Engineering a Neural Circuit
- --means: A Robust and Stable -means Variant
- Towards Automated Evaluation of Explanations in Graph Neural Networks
- A Source-Criticism Debiasing Method for GloVe Embeddings
- Visual Understanding of Multiple Attributes Learning Model of X-Ray Scattering Images
- Deep network as memory space: complexity, generalization, disentangled representation and interpretability
- Efficient Decompositional Rule Extraction for Deep Neural Networks
- Soft Autoencoder and Its Wavelet Adaptation Interpretation
- Data Summarization via Bilevel Optimization
- Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
- MUSO: Achieving Exact Machine Unlearning in Over-Parameterized Regimes
- Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models
- Fuzzy Logic Interpretation of Quadratic Networks
- Explainable Neural Computation via Stack Neural Module Networks
- Multi-Domain Transformer-Based Counterfactual Augmentation for Earnings Call Analysis
- Unified Regularity Measures for Sample-wise Learning and Generalization
- Technical Report: When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks
- Reliable Weakly Supervised Learning: Maximize Gain and Maintain Safeness
- Deep Online Learning with Stochastic Constraints
- Data Appraisal Without Data Sharing
- Bandits for Learning to Explain from Explanations
- Efficient Estimation of Influence of a Training Instance
- How far from automatically interpreting deep learning
- Accumulative Poisoning Attacks on Real-time Data
- Non-monotonic Logical Reasoning Guiding Deep Learning for Explainable Visual Question Answering
- An unsupervised framework for tracing textual sources of moral change
- Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?
- How to improve the interpretability of kernel learning
- Influence-Balanced Loss for Imbalanced Visual Classification
- Tracing the Propagation Path: A Flow Perspective of Representation Learning on Graphs
- Efficient Data-Dependent Learnability
- HERALD: An Annotation Efficient Method to Detect User Disengagement in Social Conversations
- Explainable time series tweaking via irreversible and reversible temporal transformations
- Influence Tuning: Demoting Spurious Correlations via Instance Attribution and Instance-Driven Updates
- IFBiD: Inference-Free Bias Detection
- Automated Dependence Plots
- Widen The Backdoor To Let More Attackers In
- Are You Tampering With My Data?
- Explaining Inference Queries with Bayesian Optimization
- Bandits Don't Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits
- Towards Self-Explainable Graph Neural Network
- Quantizing data for distributed learning
- Seeking the Shape of Sound: An Adaptive Framework for Learning Voice-Face Association
- On Predictive Explanation of Data Anomalies
- Adversarial Attacks on Knowledge Graph Embeddings via Instance Attribution Methods
- Longitudinal Distance: Towards Accountable Instance Attribution
- A Separation Result Between Data-oblivious and Data-aware Poisoning Attacks
- Information-theoretic Evolution of Model Agnostic Global Explanations
- Deep Active Learning by Model Interpretability
- Model-specific Data Subsampling with Influence Functions
- Wider Vision: Enriching Convolutional Neural Networks via Alignment to External Knowledge Bases
- Repairing Brain-Computer Interfaces with Fault-Based Data Acquisition
- Interactive Visual Study of Multiple Attributes Learning Model of X-Ray Scattering Images
- Energy-based Unknown Intent Detection with Data Manipulation
- DIVA: Dataset Derivative of a Learning Task
- Provable Training Set Debugging for Linear Regression
- Learning to Reweight with Deep Interactions
- Exact and Consistent Interpretation of Piecewise Linear Models Hidden behind APIs: A Closed Form Solution
- MetAL: Active Semi-Supervised Learning on Graphs via Meta Learning
- Detecting and Reducing Bias in a High Stakes Domain
- Regression Concept Vectors for Bidirectional Explanations in Histopathology
- Cross Validation for Penalized Quantile Regression with a Case-Weight Adjusted Solution Path
- Towards Sharper Utility Bounds for Differentially Private Pairwise Learning
- Explaining Anomalies in Groups with Characterizing Subspace Rules
- Model-Agnostic Explanations using Minimal Forcing Subsets
- Shapley Homology: Topological Analysis of Sample Influence for Neural Networks
- ModelPred: A Framework for Predicting Trained Model from Training Data
- AdjointBackMapV2: Precise Reconstruction of Arbitrary CNN Unit's Activation via Adjoint Operators
- Approximate Cross-Validation with Low-Rank Data in High Dimensions
- ProtoShotXAI: Using Prototypical Few-Shot Architecture for Explainable AI
- Scalable Explanation of Inferences on Large Graphs
- Logic and the -Simplicial Transformer