4 citations · 10 across the 12 of their papers we have counts for
Showing 2026 · cs.LGShow all
2 papers · 2 filters
cs.LG2026
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
Ishaan Watts, Catherine Li, Sachin Goyal +2
Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger model after subsequent changes lik…
cs.LG2026
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer +2
How can we train models whose post-trained capabilities survive subsequent fine-tuning? Rather than focusing on downstream interventions to mitigate forgetting of upstream capabili…