4 papers · 1 filter
The Implicit Bias of Logit Regularization
Alon Beck, Yohai Bar Sinai, Noam Levi
Logit regularization, the addition of a convex penalty directly in logit space, is widely used in modern classifiers, with label smoothing as a prominent example. While such method…
A Simple Model of Inference Scaling Laws
Noam Levi
Neural scaling laws have garnered significant interest due to their ability to predict model performance as a function of increasing parameters, data, and compute. In this work, we…
Probing the Latent Hierarchical Structure of Data via Diffusion Models
Antonio Sclocchi, Alessandro Favero, Noam Itzhak Levi +1
High-dimensional data must be highly structured to be learnable. Although the compositional and hierarchical nature of data is often put forward to explain learnability, quantitati…
Grokking at the Edge of Linear Separability
Alon Beck, Noam Levi, Yohai Bar-Sinai
We investigate the phenomenon of grokking -- delayed generalization accompanied by non-monotonic test loss behavior -- in a simple binary logistic classification task, for which "m…