collaborators

6 papers

cs.LG2026

Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression

Diyuan Wu, Lehan Chen, Theodor Misiakiewicz +1

It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generaliz…

cs.LG2026

Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing

Valentina Njaradi, Clémentine Dominé, Rachel Swanson +2

Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable structure from abundant unl…

cs.LG2026

Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

Peter Súkeník, Cristina López Amado, Christoph H. Lampert +1

This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions under which sinks can be represent…

stat.ML2026

High-Dimensional Private Linear Regression with Optimal Rates

Simone Bombari, Jialei Luo, Inbar Seroussi +1

Differentially private (DP) linear regression has received significant attention in the recent theoretical literature, with several approaches proposed to improve error rates. Our…

cs.LG2026

A Law of Data Reconstruction for Random Features (and Beyond)

Leonardo Iurada, Simone Bombari, Tatiana Tommasi +1

Large-scale deep learning models are known to memorize parts of the training set. In machine learning theory, memorization is often framed as interpolation or label fitting, and cl…

cs.IT2025

Contraction of Markovian Operators in Orlicz Spaces and Error Bounds for Markov Chain Monte Carlo

Amedeo Roberto Esposito, Marco Mondelli

We introduce a novel concept of convergence for Markovian processes within Orlicz spaces, extending beyond the conventional approach associated with spaces. After showing tha…