Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Learning Rate Engineering: From Coarse Single Parameter to Layered Evolution
Ming-Hong Yao, Di Wang, Jian Cui +5
Learning rate scheduling has evolved from the single global fixed rate of early SGD to sophisticated layer-wise adaptive strategies. We systematize this evolution into five generat…
cs.AI2026
Predicting LLM Output Length via Entropy-Guided Representations
Huanyi Xie, Yubin Chen, Liangyu Wang +2
The long-tailed distribution of sequence lengths in LLM serving and reinforcement learning (RL) sampling causes significant computational waste due to excessive padding in batched…