3 papers
cs.LG2025
Training Transformers with Enforced Lipschitz Constants
Laker Newhouse, R. Preston Hess, Franz Cesista +3
Neural networks are often highly sensitive to input and weight perturbations. This sensitivity has been linked to pathologies such as vulnerability to adversarial examples, diverge…
cs.LG2024
Modular Duality in Deep Learning
Jeremy Bernstein, Laker Newhouse
An old idea in optimization theory says that since the gradient is a dual vector it may not be subtracted from the weights without first being mapped to the primal space where the…
cs.LG2024
Old Optimizer, New Norm: An Anthology
Jeremy Bernstein, Laker Newhouse
Deep learning optimizers are often motivated through a mix of convex and approximate second-order theory. We select three such methods -- Adam, Shampoo and Prodigy -- and argue tha…