2 citations · 2 across the 2 of their papers we have counts for
3 papers · 1 filter
Position: Understanding LLMs Requires More Than Statistical Generalization
Patrik Reizinger, Szilvia Ujváry, Anna Mészáros +3
The last decade has seen blossoming research in deep learning theory attempting to answer, "Why does deep learning generalize?" A powerful shift in perspective precipitated this pr…
To smooth a cloud or to pin it down: Guarantees and Insights on Score Matching in Denoising Diffusion Models
Francisco Vargas, Teodora Reu, Anna Kerekes +1
Denoising diffusion models are a class of generative models which have recently achieved state-of-the-art results across many domains. Gradual noise is added to the data using a di…
Rethinking Sharpness-Aware Minimization as Variational Inference
Szilvia Ujváry, Zsigmond Telek, Anna Kerekes +2
Sharpness-aware minimization (SAM) aims to improve the generalisation of gradient-based learning by seeking out flat minima. In this work, we establish connections between SAM and…