4 papers · 1 filter
On the Asymptotics of Self-Supervised Pre-training: Two-Stage M-Estimation and Representation Symmetry
Mohammad Tinati, Stephen Tu
Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learnin…
Nearly Instance-Optimal Parameter Recovery from Many Trajectories via Hellinger Localization
Eliot Shekhtman, Yichen Zhou, Ingvar Ziemann +2
Learning from temporally-correlated data is a core facet of modern machine learning. Yet our understanding of sequential learning remains incomplete, particularly in the multi-traj…
Sharp Rates in Dependent Learning Theory: Avoiding Sample Size Deflation for the Square Loss
Ingvar Ziemann, Stephen Tu, George J. Pappas +1
In this work, we study statistical learning with dependent (-mixing) data and square loss in a hypothesis class where is the norm $\|f\|_{Î…
Shallow diffusion networks provably learn hidden low-dimensional structure
Nicholas M. Boffi, Arthur Jacot, Stephen Tu +1
Diffusion-based generative models provide a powerful framework for learning to sample from a complex target distribution. The remarkable empirical success of these models applied t…