3 papers
cs.LG2025
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
Giulia Lanzillotta, Felix Sarnthein, Gil Kur +2
The concept of knowledge distillation (KD) describes the training of a student model from a teacher model and is a widely adopted technique in deep learning. However, it is still n…
cs.LG2025
The Importance of Being Lazy: Scaling Limits of Continual Learning
Jacopo Graldi, Alessandro Breccia, Giulia Lanzillotta +2
Despite recent efforts, neural networks still struggle to learn in non-stationary environments, and our understanding of catastrophic forgetting (CF) is far from complete. In this…
cs.LG2025
Scalable Non-Equivariant 3D Molecule Generation via Rotational Alignment
Yuhui Ding, Thomas Hofmann
Equivariant diffusion models have achieved impressive performance in 3D molecule generation. These models incorporate Euclidean symmetries of 3D molecules by utilizing an SE(3)-equ…