1 paper
Yu-Hang Wu, Qin-Yuan Liu, Qiu-Yang Zhao +3
Selective layer-wise updates are essential for low-cost continued pre-training of Large Language Models (LLMs), yet determining which layers to freeze or train remains an empirical…