1 paper
Zifei Xu, Alexander Lan, Wanzin Yazar +3
Generalization abilities of well-trained large language models (LLMs) are known to scale predictably as a function of model size. In contrast to the existence of practical scaling…