4 papers
Better Training Data Attribution via Better Inverse Hessian-Vector Products
Andrew Wang, Elisa Nguyen, Runshi Yang +3
Training data attribution (TDA) provides insights into which training data is responsible for a learned model behavior. Gradient-based TDA methods such as influence functions and u…
Through a Steerable Lens: Magnifying Neural Network Interpretability via Phase-Based Extrapolation
Farzaneh Mahdisoltani, Saeed Mahdisoltani, Roger B. Grosse +1
Understanding the internal representations and decision mechanisms of deep neural networks remains a critical open challenge. While existing interpretability methods often identify…
Distributional Training Data Attribution: What do Influence Functions Sample?
Bruno Mlodozeniec, Isaac Reid, Sam Power +4
Randomness is an unavoidable part of training deep learning models, yet something that traditional training data attribution algorithms fail to rigorously account for. They ignore…
Spectral-factorized Positive-definite Curvature Learning for NN Training
Wu Lin, Felix Dangel, Runa Eschenhagen +3
Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training met…