1 paper
Taejong Joo, Wenhan Xia, Cheolmin Kim +2
Training large language models (LLMs) relies almost exclusively on dense adaptive optimizers with increasingly sophisticated preconditioners. We challenge this by showing that rand…