4 papers · 1 filter
Why Self-Training Helps and Hurts: Denoising vs. Signal Forgetting
Mingqi Wu, Archer Y. Yang, Qiang Sun
Iterative self-training (self-distillation) repeatedly refits a model on pseudo-labels generated by its own predictions. We study this procedure in overparameterized linear regress…
Training-Free Self-Correction for Multimodal Masked Diffusion Models
Yidong Ouyang, Panwen Hu, Zhengyan Wan +7
Masked diffusion models have emerged as a powerful framework for text and multimodal generation. However, their sampling procedure updates multiple tokens simultaneously and treats…
PCA++: How Uniformity Induces Robustness to Background Noise in Contrastive Learning
Mingqi Wu, Qiang Sun, Yi Yang
High-dimensional data often contain low-dimensional signals obscured by structured background noise, which limits the effectiveness of standard PCA. Motivated by contrastive learni…
Ensemble linear interpolators: The role of ensembling
Mingqi Wu, Qiang Sun
Interpolators are unstable. For example, the mininum norm least square interpolator exhibits unbounded test errors when dealing with noisy data. In this paper, we study ho…