3 papers
cs.LG2025
Revisiting Knowledge Distillation: The Hidden Role of Dataset Size
Giulia Lanzillotta, Felix Sarnthein, Gil Kur +2
The concept of knowledge distillation (KD) describes the training of a student model from a teacher model and is a widely adopted technique in deep learning. However, it is still n…
cs.LG2025
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Jonas Hübotter, Patrik Wolf, Alexander Shevchenko +3
Recent empirical studies have explored the idea of continuing to train a model at test-time for a given task, known as test-time training (TTT), and have found it to yield signific…
cs.LG2025
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
Tobias Wegel, Gil Kur, Patrick Rebeschini
We study early-stopped mirror descent (ESMD) for high-dimensional Gaussian linear regression over arbitrary convex bodies and design matrices, where the task is to minimize the in-…