8 papers
Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?
Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2
AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…
Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise
Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2
While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical…
Distributed Online Convex Optimization with Compressed Communication: Optimal Regret and Applications
Sifan Yang, Dan-Yue Li, Lijun Zhang
Distributed online convex optimization (D-OCO) is a powerful paradigm for modeling distributed scenarios with streaming data. However, the communication cost between local learners…
Continuous Subspace Optimization for Continual Learning
Quan Cheng, Yuanyu Wan, Lingyu Wu +2
Continual learning aims to learn multiple tasks sequentially while preserving prior knowledge, but faces the challenge of catastrophic forgetting when adapting to new tasks. Recent…
Non-stationary Delayed Online Convex Optimization: From Full-information to Bandit Setting
Yuanyu Wan, Chang Yao, Yitao Ma +2
Although online convex optimization (OCO) under arbitrary delays has received increasing attention recently, previous studies focus on stationary environments with the goal of mini…
Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting
Sifan Yang, Yuanyu Wan, Lijun Zhang
We investigate the online nonsubmodular optimization with delayed feedback in the bandit setting, where the loss function is -weakly DR-submodular and -weakly DR-supermodul…