2 papers
cs.LG2026
Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?
Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2
AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…
cs.LG2024
Mixture of Online and Offline Experts for Non-stationary Time Series
Zhilin Zhao, Longbing Cao, Yuanyu Wan
We consider a general and realistic scenario involving non-stationary time series, consisting of several offline intervals with different distributions within a fixed offline time…