collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2025

Revisiting Knowledge Distillation: The Hidden Role of Dataset Size

Giulia Lanzillotta, Felix Sarnthein, Gil Kur +2

The concept of knowledge distillation (KD) describes the training of a student model from a teacher model and is a widely adopted technique in deep learning. However, it is still n…

cs.LG2025

The Importance of Being Lazy: Scaling Limits of Continual Learning

Jacopo Graldi, Alessandro Breccia, Giulia Lanzillotta +2

Despite recent efforts, neural networks still struggle to learn in non-stationary environments, and our understanding of catastrophic forgetting (CF) is far from complete. In this…

cs.LG2025

Scalable Non-Equivariant 3D Molecule Generation via Rotational Alignment

Yuhui Ding, Thomas Hofmann

Equivariant diffusion models have achieved impressive performance in 3D molecule generation. These models incorporate Euclidean symmetries of 3D molecules by utilizing an SE(3)-equ…

cs.LG2024

Super Consistency of Neural Network Landscapes and Learning Rate Transfer

Lorenzo Noci, Alexandru Meterez, Thomas Hofmann +1

Recently, there has been growing evidence that if the width and depth of a neural network are scaled toward the so-called rich feature learning limit (\mup and its depth extension)…

cs.LG2024

Understanding and Minimising Outlier Features in Neural Network Training

Bobby He, Lorenzo Noci, Daniele Paliotta +2

Outlier Features (OFs) are neurons whose activation magnitudes significantly exceed the average over a neural network's (NN) width. They are well known to emerge during standard tr…