3 papers
cs.CL2024
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
Tyler A. Chang, Zhuowen Tu, Benjamin K. Bergen
How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, f…
cs.CV2024
ViTGAN: Training GANs with Vision Transformers
Kwonjoon Lee, Huiwen Chang, Lu Jiang +3
Recently, Vision Transformers (ViTs) have shown competitive performance on image recognition while requiring less vision-specific inductive biases. In this paper, we investigate if…
cs.LG2024
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
Yue Zhao, Yantao Shen, Yuanjun Xiong +5
Negative flips are errors introduced in a classification system when a legacy model is updated. Existing methods to reduce the negative flip rate (NFR) either do so at the expense…