3 papers
cs.LG2026
On the Nonlinearity of Learning Rate Scaling for LLM Training
Zaiwen Yang, Huaqing Zhang, Jing Xu +1
Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs. Exist…
cs.CL2026
QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentation
Jiazheng Li, Hongzhou Lin, Hong Lu +5
Reinforcement learning (RL) has emerged as a central paradigm for training large language models (LLMs) in reasoning tasks. Yet recent studies question RL's ability to incentivize…
cs.LG2026
Position: The Inevitable End of One-Architecture-Fits-All-Domains in Time Series Forecasting
Qinwei Ma, Jingzhe Shi, Jiahao Qiu +1
Recent work has questioned the effectiveness and robustness of neural network architectures for time series forecasting tasks. We summarize these concerns and analyze groundly thei…