2 papers
cs.LG2026
The Magic Correlations: Understanding Knowledge Transfer from Pretraining to Supervised Fine-Tuning
Simin Fan, Dimitris Paparas, Natasha Noy +3
Understanding how language model capabilities transfer from pretraining to supervised fine-tuning (SFT) is fundamental to efficient model development and data curation. In this wor…
cs.CL2025
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling
David Grangier, Simin Fan, Skyler Seto +1
Specialist language models (LMs) focus on a specific task or domain on which they often outperform generalist LMs of the same size. However, the specialist data needed to pretrain…