activity
20242026
collaborators

8 papers

cs.LG2026

Open Problem: Is AdamW Effective Under Heavy-Tailed Noise?

Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2

AdamW is the de facto optimizer for training large language models (LLMs), yet the theory behind it still lives mostly in finite-variance regimes. This is increasingly unsatisfying…

cs.LG2026

Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise

Dingzhi Yu, Hongyi Tao, Yuanyu Wan +2

While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical…

cs.LG2026

Distributed Online Convex Optimization with Compressed Communication: Optimal Regret and Applications

Sifan Yang, Dan-Yue Li, Lijun Zhang

Distributed online convex optimization (D-OCO) is a powerful paradigm for modeling distributed scenarios with streaming data. However, the communication cost between local learners…

cs.CV2025

Continuous Subspace Optimization for Continual Learning

Quan Cheng, Yuanyu Wan, Lingyu Wu +2

Continual learning aims to learn multiple tasks sequentially while preserving prior knowledge, but faces the challenge of catastrophic forgetting when adapting to new tasks. Recent…

cs.LG2025

Non-stationary Delayed Online Convex Optimization: From Full-information to Bandit Setting

Yuanyu Wan, Chang Yao, Yitao Ma +2

Although online convex optimization (OCO) under arbitrary delays has received increasing attention recently, previous studies focus on stationary environments with the goal of mini…

cs.LG2025

Online Nonsubmodular Optimization with Delayed Feedback in the Bandit Setting

Sifan Yang, Yuanyu Wan, Lijun Zhang

We investigate the online nonsubmodular optimization with delayed feedback in the bandit setting, where the loss function is -weakly DR-submodular and -weakly DR-supermodul…