3 citations · 4 across the 14 of their papers we have counts for
14 papers
Efficient Bilevel Optimization with KFAC-Based Hypergradients
Disen Liao, Felix Dangel, Yaoliang Yu
Bilevel optimization (BO) is widely applicable to many machine learning problems. Scaling BO, however, requires repeatedly computing hypergradients, which involves solving inverse…
Sketching Low-Rank Plus Diagonal Matrices
Andres Fernandez, Felix Dangel, Philipp Hennig +1
Many relevant machine learning and scientific computing tasks involve high-dimensional linear operators accessible only via costly matrix-vector products. In this context, recent a…
Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
YuXin Li, Felix Dangel, Derek Tam +1
The diagonal of a model's Fisher Information Matrix (the "Fisher diagonal") has frequently been used as a way to measure parameter sensitivity. Typically, the Fisher diagonal is es…
Kronecker-factored Approximate Curvature (KFAC) From Scratch
Felix Dangel, Bálint Mucsányi, Tobias Weber +1
Kronecker-factored approximate curvature (KFAC) is arguably one of the most prominent curvature approximations in deep learning. Its applications range from optimization to Bayesia…
Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It
Marvin F. da Silva, Felix Dangel, Sageev Oore
The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work repor…
Collapsing Taylor Mode Automatic Differentiation
Felix Dangel, Tim Siebert, Marius Zeinhofer +1
Computing partial differential equation (PDE) operators via nested backpropagation is expensive, yet popular, and severely restricts their utility for scientific machine learning.…