2 papers
stat.ML2025
Spectral-factorized Positive-definite Curvature Learning for NN Training
Wu Lin, Felix Dangel, Runa Eschenhagen +3
Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training met…
cs.LG2024
Training Data Attribution via Approximate Unrolled Differentiation
Juhan Bae, Wu Lin, Jonathan Lorraine +1
Many training data attribution (TDA) methods aim to estimate how a model's behavior would change if one or more data points were removed from the training set. Methods based on imp…