1 paper · 1 filter
Ming Li, Yanhong Li, Tianyi Zhou
What makes a difference in the post-training of LLMs? We investigate the training patterns of different layers in large language models (LLMs) through the lens of the gradient. We…