activity
20182026
most citedBackPACK: Packing more into backprop

3 citations · 4 across the 14 of their papers we have counts for

collaborators

14 papers

cs.LG2026

Efficient Bilevel Optimization with KFAC-Based Hypergradients

Disen Liao, Felix Dangel, Yaoliang Yu

Bilevel optimization (BO) is widely applicable to many machine learning problems. Scaling BO, however, requires repeatedly computing hypergradients, which involves solving inverse…

cs.LG2025

Sketching Low-Rank Plus Diagonal Matrices

Andres Fernandez, Felix Dangel, Philipp Hennig +1

Many relevant machine learning and scientific computing tasks involve high-dimensional linear operators accessible only via costly matrix-vector products. In this context, recent a…

cs.LG2025

Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator

YuXin Li, Felix Dangel, Derek Tam +1

The diagonal of a model's Fisher Information Matrix (the "Fisher diagonal") has frequently been used as a way to measure parameter sensitivity. Typically, the Fisher diagonal is es…

cs.LG2025

Kronecker-factored Approximate Curvature (KFAC) From Scratch

Felix Dangel, Bálint Mucsányi, Tobias Weber +1

Kronecker-factored approximate curvature (KFAC) is arguably one of the most prominent curvature approximations in deep learning. Its applications range from optimization to Bayesia…

cs.LG2025

Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds It

Marvin F. da Silva, Felix Dangel, Sageev Oore

The concept of sharpness has been successfully applied to traditional architectures like MLPs and CNNs to predict their generalization. For transformers, however, recent work repor…

cs.LG2025

Collapsing Taylor Mode Automatic Differentiation

Felix Dangel, Tim Siebert, Marius Zeinhofer +1

Computing partial differential equation (PDE) operators via nested backpropagation is expensive, yet popular, and severely restricts their utility for scientific machine learning.…