1 paper
Bo-Wen Zhang, Yan Yan, Boxiang Yang +2
While scaling laws optimize training configurations for large language models (LLMs) through experiments on smaller or early-stage models, they fail to predict emergent abilities d…