4 papers
Hessian Spectral Analysis at Foundation Model Scale
Diego Granziol, Khurshid Juarev
Accurate Hessian spectra of foundation models have remained out of reach, leading most prior work to rely on small models or strong structural approximations. We show that faithful…
A Linear Approach to Data Poisoning
Donald Flynn, Diego Granziol
Backdoor and data-poisoning attacks can flip predictions with tiny training corruptions, yet a sharp theory linking poisoning strength, overparameterization, and regularization is…
HessFormer: Hessians at Foundation Scale
Diego Granziol
Whilst there have been major advancements in the field of first order optimisation of deep learning models, where state of the art open source mixture of expert models go into the…
Compute-Optimal LLMs Provably Generalize Better With Scale
Marc Finzi, Sanyam Kapoor, Diego Granziol +4
Why do larger language models generalize better? To investigate this question, we develop generalization bounds on the pretraining objective of large language models (LLMs) in the…