1 paper · 2 filters
Oleg Filatov, Jan Ebert, Jiangtao Wang +1
One of the main challenges in optimal scaling of large language models (LLMs) is the prohibitive cost of hyperparameter tuning, particularly learning rate I^⋅ and batch size B.…