1 paper
Md Arafat Hossain, Xingfu Wu, Valerie Taylor +1
Despite recent advancements of large language models (LLMs), optimally predicting the model size for LLM pretraining or allocating optimal resources still remains a challenge. Seve…