2 papers
cs.LG2026
These Are Not All the Features You Are Looking For: A Fundamental Bottleneck in Supervised Pretraining
Xingyu Alice Yang, Jianyu Zhang, Léon Bottou
Transfer learning is widely used to adapt large pretrained models to new tasks with only a small amount of new data. However, a challenge persists -- the features from the original…
cs.LG2024
The Road Less Scheduled
Aaron Defazio, Xingyu Alice Yang, Harsh Mehta +3
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We pro…