1 paper
Ji Zhao, Yufei Gu, Shitong Shao +3
As Large Language Models (LLMs) achieve remarkable empirical success through scaling model and data size, pretraining has become increasingly critical yet computationally prohibiti…