6 papers
Improved Scaling Laws via Weak-to-Strong Generalization in Random Feature Ridge Regression
Diyuan Wu, Lehan Chen, Theodor Misiakiewicz +1
It is increasingly common in machine learning to use learned models to label data and then employ such data to train more capable models. The phenomenon of weak-to-strong generaliz…
Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
Valentina Njaradi, Clémentine Dominé, Rachel Swanson +2
Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable structure from abundant unl…
Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention
Peter SúkenÃk, Cristina López Amado, Christoph H. Lampert +1
This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions under which sinks can be represent…
High-Dimensional Private Linear Regression with Optimal Rates
Simone Bombari, Jialei Luo, Inbar Seroussi +1
Differentially private (DP) linear regression has received significant attention in the recent theoretical literature, with several approaches proposed to improve error rates. Our…
A Law of Data Reconstruction for Random Features (and Beyond)
Leonardo Iurada, Simone Bombari, Tatiana Tommasi +1
Large-scale deep learning models are known to memorize parts of the training set. In machine learning theory, memorization is often framed as interpolation or label fitting, and cl…
Contraction of Markovian Operators in Orlicz Spaces and Error Bounds for Markov Chain Monte Carlo
Amedeo Roberto Esposito, Marco Mondelli
We introduce a novel concept of convergence for Markovian processes within Orlicz spaces, extending beyond the conventional approach associated with spaces. After showing tha…