activity
20242026
collaborators
Showing stat.MLShow all

5 papers · 1 filter

stat.ML2026

Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning

Filip Kovačević, Hong Chang Ji, Denny Wu +2

It is folklore that reusing training data more than once can improve the statistical efficiency of gradient-based learning. While this phenomenon has been extensively studied in li…

stat.ML2025

High-dimensional Analysis of Synthetic Data Selection

Parham Rezaei, Filip Kovacevic, Francesco Locatello +1

Despite the progress in the development of generative models, their usefulness in creating synthetic data that improve prediction performance of classifiers has been put into quest…

stat.ML2025

Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery

Filip Kovačević, Yihan Zhang, Marco Mondelli

Multi-index models provide a popular framework to investigate the learnability of functions with low-dimensional structure and, also due to their connections with neural networks,…

stat.ML2025

Spurious Correlations in High Dimensional Regression: The Roles of Regularization, Simplicity Bias and Over-Parameterization

Simone Bombari, Marco Mondelli

Learning models have been shown to rely on spurious correlations between non-predictive features and the associated labels in the training data, with negative implications on robus…

stat.ML2025

High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws

M. Emrullah Ildiz, Halil Alperen Gozeten, Ege Onur Taga +2

A growing number of machine learning scenarios rely on knowledge distillation where one uses the output of a surrogate model as labels to supervise the training of a target model.…