51 citations · 95 across the 6 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.LG2023
Dual Gauss-Newton Directions for Deep Learning
Vincent Roulet, Mathieu Blondel
Inspired by Gauss-Newton-like methods, we study the benefit of leveraging the structure of deep learning objectives, namely, the composition of a convex loss function and of a nonl…
cs.LG2023
Fast, Differentiable and Sparse Top-k: a Convex Analysis Perspective
Michael E. Sander, Joan Puigcerver, Josip Djolonga +2
The top-k operator returns a sparse vector, where the non-zero values correspond to the k largest values of the input. Unfortunately, because it is a discontinuous function, it is…