4 citations · 8 across the 11 of their papers we have counts for
1 paper · 2 filters
Huan Song, Qingfei Zhao, Ting Long +4
Neural scaling laws have become foundational for optimizing large language model (LLM) training, yet they typically assume a single dense model output. This limitation effectively…