collaborators

6 papers

stat.ML2026

Understanding and Improving Shampoo and SOAP via Kullback-Leibler Minimization

Wu Lin, Scott C. Lowe, Felix Dangel +3

Shampoo and its efficient variant, SOAP, employ structured second-moment estimations and have shown strong performance for training neural networks (NNs). In practice, however, Sha…

cs.LG2026

Gauss-Newton Unlearning for the LLM Era

Lev McKinney, Anvith Thudi, Juhan Bae +4

Standard large language model training can create models that produce outputs their trainer deems unacceptable in deployment. The probability of these outputs can be reduced using…

cs.LG2025

Distributional Training Data Attribution: What do Influence Functions Sample?

Bruno Mlodozeniec, Isaac Reid, Sam Power +4

Randomness is an unavoidable part of training deep learning models, yet something that traditional training data attribution algorithms fail to rigorously account for. They ignore…

cs.LG2025

Better Training Data Attribution via Better Inverse Hessian-Vector Products

Andrew Wang, Elisa Nguyen, Runshi Yang +3

Training data attribution (TDA) provides insights into which training data is responsible for a learned model behavior. Gradient-based TDA methods such as influence functions and u…

cs.LG2025

Through a Steerable Lens: Magnifying Neural Network Interpretability via Phase-Based Extrapolation

Farzaneh Mahdisoltani, Saeed Mahdisoltani, Roger B. Grosse +1

Understanding the internal representations and decision mechanisms of deep neural networks remains a critical open challenge. While existing interpretability methods often identify…

stat.ML2025

Spectral-factorized Positive-definite Curvature Learning for NN Training

Wu Lin, Felix Dangel, Runa Eschenhagen +3

Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training met…